Image description uses a pre-trained Large Language Model (LLM) to generate textual descriptions of images and video frames.