How multimodal AI works
Classical AI systems work with one modality: a text model processes text, computer vision processes
images. Multimodal models encode data of different natures into a shared vector space, where
close vectors mean semantically similar objects — regardless of whether they started as text or as
an image.
The core principle: the string red sports jacket and a photograph of a red jacket are projected
close together, because they describe the same thing. That is what makes text-to-image and
image-to-text search possible.
Query: [photo of sneakers] → embedding: [0.82, -0.14, ..., 0.39]
Catalogue item: "Nike Air Max" → embedding: [0.79, -0.11, ..., 0.41]
Cosine similarity: 0.97 → a relevant result
Applications in e-commerce
Visual search. The shopper uploads a photo and the system finds visually similar products. It
matters most in fashion, where a similar dress is hard to put into words.
Automatic catalogue tagging. The model analyses product images and derives attributes
automatically: colour, style, pattern, visible material. That cuts the manual work of filling a
catalogue.
Description generation. Vision-language models can write product copy from photographs. Useful
for new items or for thinly filled product cards.
Enriching recommendations. Content-based filtering traditionally works with text attributes.
Adding visual embeddings makes it possible to find products that look alike, precisely where the
text characteristics are sparse or inconsistent.
Key models and technologies
| Model | Developer | Main e-commerce use |
|---|---|---|
| CLIP | OpenAI | Visual search, tagging |
| GPT-4V | OpenAI | Product descriptions, a conversational assistant with photos |
| LLaVA | Open source | Fine-tuning for specific catalogues |
| Gemini Vision | Multimodal recommendations |
Important: multimodal AI raises search latency compared with text. A photo query requires model
inference before the vector database is even queried. Caching the catalogue’s embeddings is
critical for production.