How Multimodal Models in E-commerce Learned to Understand the Product, Not the Picture
Published: 2026-10-02 · Author: AI Release · @ai_release1
⚡ The Gist in 5 Seconds - Visual search answers the question "what looks similar here?", but not "is this the same product or just a similar one?" - The idea of enriching an image with text and attributes has already been validated in Pinterest Visual Discovery, Fashion IQ, and FashionBERT. - Limitation: a single photo doesn't convey the material, composition, or many product characteristics. ### 🔍 What Was Found Vladimir Starygin, senior researcher on the retail team at tevian, breaks down the evolution of multimodal models for e-commerce. The basic visual search pipeline works like this: a product photo passes through a visual encoder and is turned into a vector, the catalog is pre-converted into a similar set of embeddings, and then it's just a matter of computing distances between vectors. In search tasks, cosine similarity is the standard choice, although Euclidean distance is also used. The closer the vectors, the higher the product ranks in the results. This scheme works as long as the needed information is actually present in both images — a condition that is far from always met. A single SKU may have five photos, a title, a description, and a dozen attributes: one shot shows the silhouette, another the sole, a third the logo, a fourth the size chart. Meanwhile, the attributes may specify a material that can't be determined from the photo. Two white T-shirts — regular cotton and athletic synthetic — will look very similar visually with the same angle and a similar cut. Back in 2017, Pinterest described its Visual Discovery system, where convolutional network features were used to search for visually similar images and objects, while Pinterest Lens allowed users to select an object in a photo and search for similar items within the catalog. Then came Fashion IQ with its interactive fashion retrieval task: the user explains in words how the desired item should differ from the original one, and refines the query with new phrasing if the wrong image is returned. A typical query would be — "I like this dress, but I want something similar with long sleeves." The core architecture is a multimodal transformer that combines image features, product attributes, the user's textual feedback, and the history of previous dialogue steps. ### 💡 Why It Matters The key question driving the multimodal side of e-commerce is: what exactly should be packed into the embedding — a single photo, or a representation of the product itself assembled from all available information. A product already has several sources: photos show shape, color, and visual details; the title may contain the model and purpose; attributes cover size, material, and composition; sometimes user behavior is added on top. For finding a specific SKU, the distinction is fundamental: two sneaker models from the same brand with the same color and a nearly identical silhouette, but a different sole and a small insert on the heel — for a "something similar" query, both results are fine, but for an exact search, one of them is already a mistake. ### 🧩 Context An intermediate step before CLIP was FashionBERT: in 2020, even before the era
⚡ The Gist in 5 Seconds - Visual search answers the question "what looks similar here?", but not "is this the same product or just a similar one?" - The idea of enriching an image with text and attributes has already been validated in Pinterest Visual Discovery, Fashion IQ, and FashionBERT.
- Limitation: a single photo doesn't convey the material, composition, or many product characteristics.
🔍 What Was Found Vladimir Starygin, senior researcher on the retail team at tevian, breaks down the evolution of multimodal models for e-commerce.
The basic visual search pipeline works like this: a product photo passes through a visual encoder and is turned into a vector, the catalog is pre-converted into a similar set of embeddings, and then it's just a matter of computing distances between vectors.
In search tasks, cosine similarity is the standard choice, although Euclidean distance is also used.
The closer the vectors, the higher the product ranks in the results.
This scheme works as long as the needed information is actually present in both images — a condition that is far from always met.
A single SKU may have five photos, a title, a description, and a dozen attributes: one shot shows the silhouette, another the sole, a third the logo, a fourth the size chart.