All terms

Consumer Behaviour

Multimodal search

Also known as: multi-modal search, multimodal query

Multimodal search is search that combines more than one type of input in a single query, such as an image plus text, or voice plus a screenshot. Instead of describing something in words alone, the user shows the system what they mean and adds a question or refinement. Search engines and assistants interpret the combined signals to return one set of results.

What it is

Multimodal search lets a person query using a mix of images, text, voice and sometimes video or screen content at the same time. A typical example is photographing a chair and typing "in green velvet" to find variants. The underlying models are trained to relate visual and language inputs to the same set of concepts, so the two parts of the query are treated as one intent.

Why it matters

It changes how demand reaches your site, because the entry point may be a photograph of your product, a competitor's product or a physical object with no brand name attached. Queries that were previously impossible to express in words, such as "what is this part called", now convert into commercial intent. If your images, product data and page copy do not describe an item the way a model would recognise it, you can be excluded from results you would otherwise win.

How it works

The system encodes the image and the text into a shared representation, retrieves candidate matches, then uses the text portion to filter or modify the visual match. Practitioners prepare for this by publishing clear, well lit, uncluttered product photography from several angles, writing descriptive alt text and captions, keeping structured product data accurate, and making sure attributes such as colour, material, size and model number appear as text on the page. Testing your own catalogue through Lens style tools shows which products are recognised and which are confused with others.

When it applies

It applies most in retail, home improvement, fashion, parts and components, plants and food, travel and any category where the object is easier to show than to name. It also applies to on screen research, where a user queries something they are already looking at rather than starting a new search.

Examples

  • A shopper photographs a pair of trainers in the street and adds "in size 9 under £90" to find stock nearby.
  • A homeowner points a camera at a boiler control panel and asks "how do I reset this", landing on a manufacturer support page.
  • A designer screenshots a lamp from a social post and adds "similar in brass" to find comparable products from other retailers.

How it is measured

  • Share of image or Lens driven referrals in analytics, where the source is identifiable
  • Product coverage rate: proportion of catalogue items correctly recognised in manual camera tests
  • Image impressions and clicks in Google Search Console, split by page type
  • Conversion rate and average order value for sessions that begin on image heavy landing pages

Related terms in Consumer Behaviour

Primary research · August 2026

How ChatGPT Shortlists Software Brands

An audit across 10 categories and 60 buying questions. I recorded what ChatGPT reads, throws away and links to when a buyer asks it which software to buy, and what that decides.

60
Questions asked
10
Software markets
2,680
Results read
367
Links shown
Free35 pages · PDF · 536 KBDiscovery Digest every Friday

Free download

Get the full report

35 pages · PDF · 536 KB. Enter your details and it downloads straight away.

How ChatGPT Shortlists Software Brands downloads straight away. No spam, unsubscribe anytime.