What a vector database is

A vector database stores objects as numeric vectors — embeddings — and specialises in one job:
finding the most similar objects to a query vector, quickly.

The standard flow in e-commerce:

Product → embedding model → vector [0.23, -0.41, 0.87, ...] → vector database
Query   → embedding model → vector [0.21, -0.39, 0.85, ...] → ANN search → top 20 products

Cosine distance between vectors is the measure of semantic closeness. Two vectors near each other in
the space mean two objects similar in meaning, regardless of how they are worded.

Why a specialised database is needed

A relational database is a poor fit for similarity search in high-dimensional spaces. At a million
products and 768-dimensional vectors, an exact scan on every query would take seconds. Vector
databases use dedicated indexes:

Index Principle When it fits
HNSW A graph of nearest neighbours High accuracy, moderate volume
IVF Clustering plus in-cluster search Very large catalogues
PQ Vector quantisation Memory savings

HNSW (hierarchical navigable small world) is the most widely used: under 10 ms lookup at recall
above 95%.

Applications in e-commerce

Semantic search. A query such as a warm jacket for mountain hiking becomes an embedding and is
matched against the catalogue of product descriptions. Relevant results appear even with no keyword
overlap.

Similar items recommendations. The embedding of the viewed item drives a nearest-neighbour
lookup in product space. It works on cold items with no interaction history.

AI shopping assistants and RAG. An LLM assistant receives context through retrieval — a vector
search across descriptions, specifications and product FAQs.

Tip: for hybrid search (vector relevance plus keywords plus filters), choose a database with
native hybrid support, or pair pgvector with PostgreSQL full-text search.

Common implementations

  • Qdrant — Rust, high performance, open source, a good fit for self-hosted deployment
  • Pinecone — a managed service, simple integration, serverless tier
  • Weaviate — hybrid search, modular architecture
  • Milvus — scales to very large collections
  • pgvector — a PostgreSQL extension, convenient when the infrastructure is already on Postgres