Back to insights
A magnifying glass resting on a black and white photograph of an acacia tree on open plain, with the tree beneath the glass redrawn as a gold diagram of fine lines running out to small blank labels.

What AI Search Does With Your Images

Google Lens handles close to 20 billion visual searches a month. That is Google’s own number, and every one of those searches contains no words at all. Someone pointed a camera at a thing and asked about it.

Most writing about AI and search is about text: what happens to your article when a summary answers the question instead. That is worth understanding, and we went through the evidence. This is the other half, and almost nobody covers it. Your pictures are now being read, by more things, in more ways, and the practices that matter are the boring ones you already knew about.

A search with no words in it

Visual search inverts the usual order. Instead of describing a thing in order to find a picture of it, someone shows the picture and asks what the thing is, where to buy it, what it costs. The query never existed as text, so there was never a keyword to rank for.

What decides whether your image is the answer is whether the system can work out what it shows. Which means the question stops being “what did they type” and becomes “what does this picture demonstrably contain, and can a machine tell”.

What actually gets read

Google is direct about this for image search. It says it uses alt text “along with computer vision algorithms and the contents of the page to understand the subject matter of the image”. Three inputs, weighted together, none of them sufficient alone.

The models answering questions in AI summaries and chat assistants work the same way and add to it. They are multimodal, which means the image itself is an input rather than something referred to. So for one photograph on one page, the things being read are:

  • The pixels. What is visibly in the frame, read directly.
  • The filename. Cheap to read, available before anything is decoded.
  • The alt text. Your own description of what the image shows and why it is here.
  • The caption and the surrounding paragraph. What the page claims about it.

The interesting part is that these can disagree, and a disagreement is expensive. A photograph of a kitchen, in a file called property-4.jpg, with alt text reading “house”, on a page about a listing in Claremont, gives four weak and partly conflicting accounts of one picture. A system reconciling those has no reason to be confident about any of them.

Alt text got promoted

For twenty years alt text was an accessibility obligation that quietly also helped search. It was the thing you added last, or did not add, and the cost of skipping it was borne by people using screen readers, which is precisely why it kept getting skipped.

It is now also the sentence that tells a machine what your picture is, in its own words, in a form no vision model has to infer. That has not changed what good alt text looks like. It has changed how much of your work depends on it.

The advice is the same advice, which should be reassuring rather than disappointing: describe what is there and why it matters on this page, concretely, without keyword stuffing. Twenty-five before and after rewrites has the working version, and the rules behind it come from the accessibility standard rather than from anyone’s theory about ranking.

Filenames are still the cheapest thing you can fix

A filename is read before an image is decoded, costs nothing, and is attached to the file wherever it travels. It leaves the page with the image when someone downloads it, and it is what you search your own drive for in two years.

Google describes a filename as giving “very light clues” about an image’s subject. Light is not nothing, and it is the same nothing whether one system is reading it or four are. DSC_0042.JPG tells every one of them the same thing, which is that you did not say.

Weight matters more, not less

It would be easy to assume that once machines are doing the reading, how heavy an image is stops mattering. The opposite is true, for a plain reason: the people are still there. A slow page is still slow, and images are still the heaviest thing on almost every page.

What has changed is that there is now a second reason. Retrieval systems crawl at scale and under budget. A page that takes four seconds to deliver its images is a page that is more expensive to read, and it is read by more things than it used to be. The format comparison has the measurements, and the short version is that AVIF was 57 to 58% smaller than JPEG on our own files.

What to actually do

None of this is new work. It is the same list, with more riding on it.

  • Make the four signals agree. Filename, alt text, caption and surrounding text should describe the same picture. This is the highest-value thing on the list and the one most often wrong.
  • Describe, do not label. “Two lions resting in dry grass on a game reserve” is worth more than “lions” to a reader and to a model, because it is falsifiable against the pixels.
  • Name the file for the subject, in plain words, hyphen separated, before it reaches your CMS. Renaming afterwards means a URL change and redirects.
  • Keep it small. Resize to the size it is displayed at and serve a modern format.
  • Do it on every image, not the memorable ones. Consistency across a library is worth more than excellence on six photographs, and it is the part that fails under deadline.

Questions people ask

Does this mean image SEO is more important now?

It means the same practices are read by more systems. Whether that is “more important” depends on your site. For a shop whose customers point cameras at products, considerably. For a text-heavy site with decorative photography, not really.

Should I write alt text for the machine or the person?

The person. This is not a compromise: a description that genuinely helps someone who cannot see the image is also the most useful description a model can read, because both want the same thing, which is an accurate account of what is there. Writing for the machine is how you end up with keyword-stuffed alt text that Google warns about by name.

Can I not just let the AI work it out from the picture?

It will try, and it will often be right about what is in the frame and wrong about why it matters. A model can see a bedroom. It cannot see that this is the Garden Suite at a particular lodge, or that the fabric is the thing the customer is deciding about. That context only exists if you supply it.

Is there a shortcut?

Doing it in bulk, at the point the images enter your site, rather than one file at a time afterwards. Resiqo names a whole folder, writes the alt text for each one and resizes and converts them in the same pass, in your browser, so the originals never leave your machine. The work still has to happen. It does not have to happen by hand.


Sources: Google, on Lens handling nearly 20 billion visual searches a month; Google Search Central, Image SEO best practices, for how alt text, computer vision and page content are used together, and for the “very light clues” description of filenames. The file size comparisons referred to above were measured on this site’s own images and are set out in full in the formats article.