Customer Impact

SEO & GEO

Image Alt Text: Optimising Images for Multimodal AI

Copy for AI

Multimodal AI can now look at an image itself, but image alt text still matters because the model uses that text to understand what a visual means in the context of your page. A model sees the pixels, but it reads your alt text, caption and surrounding copy to work out why that image is there and which topic it belongs to. In this article you will learn how multimodal models process images, how to write alt text and captions that work for AI, and how to link your images to the right entity with ImageObject markup.

How do multimodal AI models read an image?

A multimodal model processes text and images in the same meaning space, so it can compare a picture with words. Models such as GPT-4o, Gemini and Claude can interpret an image directly: they recognise objects, read text inside the image and summarise what can be seen. Under the hood, they turn both your text and your image into vectors, the same way AI turns meaning into numbers. That is how a model can match “a chart showing rising revenue” to a question about growth, even when those words do not literally appear in the image.

But there is a limit to what a model can extract from pixels alone. A photo of two people at a table could be a sales conversation, a job interview or a coffee break. The image on its own is ambiguous. The alt text, the caption and the paragraph around it remove that ambiguity. They tell the model not only what is shown, but why it is shown and which topic it belongs to. For AI, an image is therefore rarely a standalone object: it is an element that takes its meaning from the text around it.

Does alt text still matter if AI can see images?

Yes, alt text still matters, because it remains the most direct and reliable way to hand the meaning of an image to a machine. Three reasons make that clear.

First, accessibility: alt text is written for people using a screen reader, and that function is not going away. In Europe, accessible web content is also increasingly a legal requirement, so this is not an optional extra in any case.

Second, reliability: image recognition is good, but it is neither perfect nor always consistent. Alt text is an explicit, unambiguous signal that you control yourself. You do not leave the interpretation to the model, you steer it.

Third, coverage: not every system that processes your content looks equally deeply at your visuals. Classic search engine crawlers, AI systems that process large volumes of pages quickly, and assistive technology all still lean heavily on text. The annual WebAIM analysis of highly visited home pages shows that roughly one in six images has no alt text. That is not a small detail: it means those visuals literally say nothing to a large part of the web.

How do you write alt text that works for AI?

Good alt text describes the function and meaning of an image within your story, not a string of loose keywords. The old reflex of stuffing “alt text seo” and five variants into every alt attribute backfires: the model reads meaning, and a list of keywords reads as noise. Write a short, complete sentence that fits the context instead.

A few concrete guidelines:

  • Describe what is happening, not just what is there. “Dashboard showing the AI visibility score per month” is stronger than “dashboard screenshot”.
  • Stay aligned with the topic of the page. If the paragraph is about lead follow-up, let the alt text show that connection. The image then reinforces the meaning of your copy instead of sitting beside it.
  • Keep it short and natural. One sentence or half a sentence is enough. Write the way you would explain to a colleague what the image shows.
  • No keyword stuffing. One natural mention of the topic is enough. Repetition narrows your vector, it does not strengthen it.
  • Skip purely decorative visuals. A mood line or background shape gets an empty alt (alt=""), so a screen reader and a crawler know there is no content to miss.

The core question is simple: if someone could not see the image, which sentence would best replace the information? That is your alt text.

What is the difference between alt text and a caption?

Alt text is a hidden description for machines and screen readers, while a caption is visible text that every reader sees below or next to the image. They serve a different purpose, and AI uses both.

Alt text replaces the image when it cannot be loaded or seen. A caption adds context for anyone who does see the image: a source, an explanation, a conclusion. For multimodal AI, a caption is especially valuable, because it often carries the interpretation of the image that you want to convey yourself. A chart captioned “AI referrals rose after restructuring the product pages” tells the model right away which story the image supports.

In practice, they reinforce each other. The alt text describes the image factually, the caption gives it meaning. Together they make an image extractable: an AI can pull a small, citable piece of information out of it. That fits the broader logic of content architecture for AI extraction, where every element of your page is independently readable and summarisable.

What does ImageObject markup do for your images?

ImageObject markup explicitly links an image to an entity and gives AI structured facts about that visual, separate from the HTML around it. It is a piece of schema.org data, usually in JSON-LD, in which you provide fields such as contentUrl (where the image lives), caption, description and, for rights, creator, creditText and license.

The value lies in certainty. In plain HTML, a system has to infer which image a piece of text belongs to. With ImageObject you make that connection unambiguous: this visual belongs to this description and this source. For AI systems that determine which source is accurate and what they may cite, such an explicit link is exactly the kind of anchor that increases the chance of a correct mention.

A few practical points:

  • Reserve markup for images that genuinely matter. A main product photo, a data visualisation or an author portrait deserves ImageObject. Marking up every decorative icon is pointless.
  • Let caption and description complement each other. The caption is short and concrete, the description can provide the context.
  • Keep the data consistent with what is visible on the page. Markup that deviates from the actual content undermines your reliability instead of strengthening it.
  • Choose an efficient format. Modern formats such as WebP and AVIF load faster, and speed remains a factor in how well your images are processed and displayed.

The short summary

Multimodal AI can see your images, but it only really understands them through the text around them. Alt text gives an unambiguous description, a caption adds meaning, and ImageObject markup ties the whole thing to the right entity. So treat your images as full-fledged content, not decoration: describe their function, stay aligned with your topic and avoid keyword stuffing. That is how your visuals become findable and citable instead of mute.

For a B2B company, what ultimately counts is not how beautiful an image is, but whether it contributes to visibility that produces leads. Want to know how to make your images and the rest of your content work in AI search engines? Explore our approach to generative engine optimization, or start with the basics in our GEO guide.

Book your free intake

Free website scan

Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.

Where should we send your report?

We only use your details for your scan. No spam, unsubscribe anytime.