Vector Database

A vector database is a database system specifically designed for storing, indexing, and retrieving high-dimensional vector data.

You can think of it as: storing things with similar meanings together, and quickly finding the ones most similar to a given item.

Unlike traditional databases, which query through exact matching (WHERE name = 'Alice'), vector databases query through similarity (finding the 10 images most similar to this image).

An intuitive analogy

Imagine a library scenario:

Database typeRetrieval methodAnalogy
Traditional databaseExact search by book number and titleFind a book with a specified number
Vector databaseSearch by content relevanceFind "all science fiction novels similar in style to The Three-Body Problem"

This semantic similarity is precisely the core problem that vector databases solve.


Why do we need a vector database?

Before diving into technical details, let's first understand what problems vector databases solve.

Limitations of traditional databases

Traditional relational databases (MySQL, PostgreSQL) are very good at handling structured data, but they fall short when faced with the following requirements:

  • Image search (find visually similar images)
  • Semantic search (when a user searches for "Apple phone", it can find related content about "iPhone")
  • Recommendation systems (find "songs similar in style to the ones you like")
  • Anomaly detection (find "logs that differ the most from normal behavior")

The common characteristic of these problems is the need to understand the "meaning" of content, rather than performing literal matching.

Problems with traditional approaches

用 LIKE '%苹果%' 搜索 → 找不到 "iPhone"、"Apple"
用全文索引搜索     → 找不到语义相关但用词不同的内容

Comparison diagram

The diagram below intuitively shows the fundamental differences in query methods between traditional databases and vector databases.

Traditional database vs. vector database: comparison of query methods Traditional database (exact match) SELECT * WHERE name = 'Apple phone' Match results: √ Apple phone Pro 128GB x iPhone 15 (no match) x Apple mobile phone (no match) x smartphone iOS (no match) Vector database (semantic similarity) search(embed("iPhone"), top_k=4) Similar results (with similarity scores): √ Apple phone Pro 128GB 0.98 √ iPhone 15 0.95 √ Apple mobile phone 0.93 √ smartphone iOS 0.87

Core concepts: vectors and embeddings

Understanding vectors and embeddings is the first step to mastering vector databases.

What is a vector?

In mathematics, a vector is an ordered set of numbers.

[0.12, -0.54, 0.87, 0.03, ..., 0.61]   ← 这就是一个向量

In machine learning, this set of numbers represents the semantic features of an object, typically with dimensions ranging from 128 to 4096.

What is an embedding?

Embedding is the process and result of converting real-world objects (text, images, audio, etc.) into vectors.

This conversion is performed by an embedding model, whose core idea is: objects with similar semantics have vectors that are closer in space.

Embedding process illustration Text "The weather is really nice today" Image A photo of a cat Audio A music clip Embedding model Embedding Model text-embedding-3 CLIP / ResNet ... Text vector (1536 dimensions): [0.12, -0.54, 0.87, 0.03, ...] Image vector (512 dimensions): [-0.33, 0.71, 0.22, 0.95, ...] Audio vector (256 dimensions): [0.66, -0.11, 0.48, -0.72, ...] Tip: Objects with similar semantics have vectors that are closer in space after conversion.

Similar semantics means similar vectors

Use a simplified 2D example to understand (in reality it is hundreds to thousands of dimensions):

Semantic clustering in vector space (2D illustration) x y Animals cat dog Rabbit bear Technology Computer Phone Keyboard Monitor Food Pizza Hamburger Noodles Query: pet The vector for the query "pet" is closer to "cat" and "dog", belonging to the animal cluster.

Key understanding: Two vectors that are close in vector space also have semantically similar original content. This is the foundation of all capabilities of vector databases.


Similarity calculation methods

The core of finding the "most similar vector" is to calculate the distance or similarity between two vectors. Here are the three most commonly used methods.

Cosine Similarity

Cosine similarity measures the angle between two vectors, ignoring their magnitude. It is the most commonly used method, especially suitable for text scenarios.

Formula:

\[ \text{CosineSimilarity}(A, B) = \frac{A \cdot B} {\|A\|\|B\|} = \frac{\sum_{i=1}^{n} A_i B_i} {\sqrt{\sum_{i=1}^{n} A_i^2}\sqrt{\sum_{i=1}^{n} B_i^2}} \]

  • Result range: -1 to 1, larger value means more similar
  • Applicable scenarios: text semantic search, document similarity

Euclidean Distance

Euclidean distance measures the straight-line distance between two points; the smaller the distance, the more similar they are.

Formula:

\[ d(A,B) = \sqrt{\sum_{i=1}^{n}(A_i-B_i)^2} \]

  • Result range: 0 to ∞, smaller value means more similar
  • Applicable scenarios: image retrieval, location-related applications

Dot Product

The dot product is the sum of vector component products, combining both direction and magnitude information.

Formula:

\[ A \cdot B = \sum_{i=1}^{n} A_i B_i \]

  • Applicable scenarios: recommender systems (equivalent to cosine similarity when vectors are normalized)

Comparison of the three methods

Comparison of three similarity calculation methods Method Principle Meaning of the result Recommended scenarios Cosine similarity Compute the cosine value of the angle between two vectors Focuses on direction, ignores magnitude [-1, 1], the closer to 1, the more similar Text search, first choice for NLP Euclidean distance The straight-line distance between two points Focuses on absolute position differences [0, ∞), the closer to 0, the more similar Image retrieval, coordinate system data Dot product Sum of the products of the vector components Comprehensive consideration of direction + magnitude The larger the value, the more similar (no fixed range) Recommender systems, normalized scenarios

Python code example

The following example demonstrates Python implementations of three similarity calculation methods:

Example

import numpy as np

# Cosine similarity: measures directional similarity
def cosine_similarity(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

# Euclidean distance: measures absolute position differences
def euclidean_distance(a, b):
    return np.linalg.norm(a - b)

# Dot product: combines direction and magnitude
def dot_product(a, b):
    return np.dot(a, b)

# Example vectors
v1 = np.array([0.12, -0.54, 0.87, 0.03])
v2 = np.array([0.10, -0.50, 0.90, 0.05])
v3 = np.array([-0.80, 0.20, -0.30, 0.70])

print(f"v1 vs v2 cosine similarity: {cosine_similarity(v1, v2):.4f}")  # Approximately 0.9997 (very similar)
print(fv1 vs v3 cosine similarity: {cosine_similarity(v1, v3):.4f})  # Approximately -0.55 (not similar)
v1 vs v2 余弦相似度: 0.9997
v1 vs v3 余弦相似度: -0.5512

Vector indexing algorithms

When the data volume is large (millions, billions), computing similarity for every piece of data (brute-force search) is too slow. Vector databases use specialized indexing algorithms to accelerate queries.

Brute-force search (Flat / Brute-force)

Brute-force search traverses all vectors and computes similarity one by one.

DimensionDescription
PrincipleTraverse all vectors, compute similarity one by one
Advantages100% accurate results
DisadvantagesExtremely slow with large data volumes, O(n) complexity
ApplicableData volume less than 100k, extremely high precision requirements

IVF (Inverted File Index)

IVF index principle: first cluster, then search within buckets Cluster 1 center Cluster 2 center Cluster 3 center Query far Close! far Perform exact search only within the nearest cluster 2, greatly reducing computation

IVF execution steps:

  1. Training phase: Use K-Means to cluster all vectors into N clusters, record the center of each cluster
  2. Query phase: First find the centers of the nearest clusters, then perform exact search only within these clusters

HNSW (Hierarchical Navigable Small World graph)

HNSW is currently the most mainstream vector indexing algorithm, balancing speed and accuracy.

HNSW hierarchical structure diagram Layer 2 (sparsest, long-range jumps) A B Layer 1 (medium density, medium-range jumps) C A D B E Layer 0 (densest, exact search) F C A G D H B E I At query time, start from the top layer with large jumps to locate the region, then refine layer by layer to precisely find the nearest neighbor

HNSW core idea:

  • Build a multi-layer graph structure, sparse at the top, dense at the bottom
  • At query time, start from the top-layer entry point and play a "hopscotch game": greedily jump to closer nodes at each layer, then descend to the next layer
  • Greatly reduces the number of nodes that need to be compared, time complexity approximately O(log n)

Other common indexes

Index typeFeaturesApplicable scenarios
Flat (brute force)Accurate but slowSmall datasets, accuracy first
IVF_FlatExact search after clustering, fastMedium to large scale, sufficient memory
IVF_PQQuantization compression, saves memoryUltra-large scale, memory constrained
HNSWFast, high accuracy, high memory usageMost commonly used, recommended first choice
ScaNNMade by Google, optimized throughputHigh-concurrency production environments

Comparison of mainstream vector databases

The following is a horizontal comparison of the most mainstream vector databases currently available, helping you make choices in different scenarios.

Horizontal comparison of mainstream vector databases Database Type Deployment method Features Applicable scenarios Difficulty Chroma Open source and free Pure vector DB Local / Cloud Embedded-first Minimalist API, Python-native Easiest to integrate with LangChain RAG prototyping, AI application development Beginner Qdrant Open source and free Pure vector DB Local / Docker Cloud service Rust implementation, high performance Supports hybrid filtering + vector search Production-grade recommendation, performance first Intermediate Weaviate Open source and free Multimodal DB Local / Cloud SaaS GraphQL API, built-in vectorization Multimodal (text + image) Multimodal retrieval, knowledge graph Intermediate Milvus Open source and free Pure vector DB Distributed deployment Kubernetes LF AI Foundation project Large-scale distributed, comprehensive features Hundreds of millions of data, enterprise scale Advanced Pinecone Commercial SaaS Managed vector DB Purely cloud-based Fully managed service Zero ops, out-of-the-box Free tier available Quick launch, for teams without ops capabilities Beginner pgvector Open source plugin PG extension Existing PostgreSQL Use directly in your environment Reuse existing PG infrastructure SQL interface, fastest to learn Projects already using PG, lightweight integration Beginner

Beginner advice: Start with Chroma or pgvector. The former is suitable for AI application prototypes, the latter for projects with existing PostgreSQL.


Quick start: Python examples

Below we use Chroma (easiest to get started) to demonstrate the complete CRUD workflow.

Installation

Example

pip install chromadb openai

Complete example: building a document semantic search system

The following code demonstrates end-to-end how to use Chroma to build a semantic-based document search system.

Example

import chromadb
from chromadb.utils import embedding_functions

# ─── 1. Initialize client ───────────────────────────────────────────
# Persist locally (recommended)
client = chromadb.PersistentClient(path="./my_vector_db")

# Use OpenAI embedding model (could also switch to a local model)
openai_ef = embedding_functions.OpenAIEmbeddingFunction(
    api_key="your-openai-api-key",       # Required: replace with your API Key
    model_name="text-embedding-3-small"   # 1536 dimensions, cost-effective
)

# ─── 2. Create collection (like a "table" in relational databases)────────────────────────
collection = client.get_or_create_collection(
    name="my_documents",                   # Collection name
    embedding_function=openai_ef,          # Bind embedding function
    metadata={"hnsw:space": "cosine"}      # Use cosine similarity
)

# ─── 3. Insert Documents ──────────────────────────────────────────────
documents = [
    Python is an object-oriented interpreted programming language, widely used in data science and AI development,
    Machine learning is a subfield of artificial intelligence that enables computers to learn patterns from data,
    Deep learning uses multi-layer neural networks and performs excellently in image recognition and NLP tasks,
    Vector databases are specifically designed to store high-dimensional vectors and support semantic similarity search,
    PostgreSQL is a powerful open-source relational database,
    Redis is an in-memory high-performance key-value database, often used for caching,
    Docker containerization technology allows applications to run consistently in any environment,
    Git is a distributed version control system and a fundamental tool for modern software development,
]

ids = [f"doc_{i}" for i in range(len(documents))]

# Batch insert (Chroma automatically calls the embedding model to convert to vectors before storing)
collection.add(
    documents=documents,
    ids=ids,
    metadatas=[{"source": "tutorial", "index": i} for i in range(len(documents))]
)

print(fInserted {len(documents)} documents)
已插入 8 条文档

Example

# ─── 4. Semantic Search ──────────────────────────────────────────────
query = How to do artificial intelligence with Python

results = collection.query(
    query_texts=[query],
    n_results=3,                # Return the top 3 most similar
    include=["documents", "distances", "metadatas"]
)

print(f"\n"Query: {query}")
print("-" * 50)
for i, (doc, dist) in enumerate(zip(
    results["documents"][0],
    results["distances"][0]
)):
    similarity = 1 - dist   # Convert cosine distance to similarity
    print(fRank {i+1} (similarity {similarity:.4f}):)
    print(f"  {doc}")
    print()
查询:如何用 Python 做人工智能
--------------------------------------------------
第 1 名(相似度 0.9231):Python 是一种面向对象的解释型编程语言...
第 2 名(相似度 0.8874):机器学习是人工智能的子领域...
第 3 名(相似度 0.8612):深度学习使用多层神经网络...

Example

# ─── 5. Search with Filters (Metadata Filtering) ───────────────────────
results_filtered = collection.query(
    query_texts=[Database technology],
    n_results=2,
    where={"source": "tutorial"},        # Search only among documents where source=tutorial
    include=["documents", "distances"]
)

# ─── 6. Update Documents ──────────────────────────────────────────────
collection.update(
    ids=["doc_0"],
    documents=[Python is currently the most popular programming language, widely used in AI, data analysis, and web development],
    metadatas=[{"source": "tutorial", "index": 0, "updated": True}]
)

# ─── 7. Delete Documents ──────────────────────────────────────────────
collection.delete(ids=["doc_7"])   # Delete Git-related documents

# ─── 8. View Collection Statistics ──────────────────────────────────────────
print(fCurrent collection document count: {collection.count()})

Without using third-party embedding APIs (fully local)

If you don't want to use the OpenAI API, you can run completely offline with local embedding models.

Example

import chromadb
from sentence_transformers import SentenceTransformer

# Use local embedding models (no API key required, fully offline)
model = SentenceTransformer("paraphrase-multilingual-MiniLM-L12-v2")  # Supports Chinese

client = chromadb.Client()
collection = client.create_collection("local_demo")

texts = ["The weather is nice today", "It's sunny and perfect for going out", "The stock market surged", "It might rain tomorrow"]

# Manually generate vectors before insertion
embeddings = model.encode(texts).tolist()
collection.add(
    embeddings=embeddings,
    documents=texts,
    ids=[f"id_{i}" for i in range(len(texts))]
)

# Query
query_embedding = model.encode(["What's the weather like today?"]).tolist()
results = collection.query(query_embeddings=query_embedding, n_results=2)
print(results["documents"])
# Output: [['The weather is nice today', 'It's sunny and perfect for going out']]

pgvector example (for PostgreSQL users)

If your project already uses PostgreSQL, pgvector is the lightest way to integrate.

Example

-- Install the extension
CREATE EXTENSION IF NOT EXISTS vector;

-- Create table: store article titles and their vectors (1536 dimensions)
CREATE TABLE articles (
    id       SERIAL PRIMARY KEY,
    title    TEXT NOT NULL,
    content  TEXT,
    embedding vector(1536)           -- Vector column, 1536 dimensions
);

-- Create HNSW index to speed up queries
CREATE INDEX ON articles
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);

-- Insert data (vectors are generated in the application layer and passed in)
INSERT INTO articles (title, embedding)
VALUES ('Python Beginner's Guide', '[0.12, -0.54, 0.87, ...]'::vector);

-- Semantic search: find the 5 most similar articles
SELECT id, title,
       1 - (embedding <=> '[0.10, -0.50, 0.90, ...]'::vector) AS similarity
FROM articles
ORDER BY embedding <=> '[0.10, -0.50, 0.90, ...]'::vector
LIMIT 5;

Note: <=> is a vector distance operator provided by pgvector, used to compute cosine distance. Subtract the cosine distance from 1 to get cosine similarity.


Typical application scenarios

Vector databases have a wide range of applications in the AI era; here's an overview of the six most typical scenarios.

Typical application scenarios of vector databases RAG RAG knowledge base Q&A Convert enterprise documents into vectors, When users ask questions, retrieve relevant fragments Hand over to LLM to generate answers. Examples: ChatPDF, NotionAI Enterprise knowledge base assistant Recommendation Personalized recommendation systems Convert user behavior and product information into vectors, find "users with similar tastes what they like" for recommendations. Examples: Spotify song recommendations E-commerce related product recommendations Image Search by image Encode images into vectors, Find visually similar images through similarity search, no labels needed. Examples: Google Image Search Taobao photo search for products Detection Anomaly detection Normal behavior is mapped to clustered Vector region, vector of abnormal behavior. Move away from the normal area to detect it. Representative: Network Intrusion Detection Financial Fraud Detection deduplication Content deduplication / copyright detection Determine two passages based on vector similarity. Identify whether the content is highly similar Plagiarized or duplicate content. Representative: Paper Plagiarism Detection System Music copyright detection face Face / biometric recognition Encode the face as a feature vector, Retrieve the most similar in the vector database Known face recognition completed. Representative: Face Access Control System Mobile face unlock

RAG (Retrieval-Augmented Generation) architecture

RAG is currently one of the primary application scenarios for vector databases. Below is its core workflow:

RAG system workflow Step 1 User question "How does Python handle exceptions?" Step 2 Question vectorization. Embedding model -> [0.12, -0.54, ...] Step 3 Vector retrieval In vector DB Find most similar documents Step 4 LLM generation Retrieve content + question -> Prompt -> Answer Step 5 Return answer Document-based. Accurate, traceable answers The vector database plays a core role in step 3: finding the most relevant knowledge snippets in milliseconds.

Selection recommendations and best practices

Selection decision tree

Based on your specific situation, choose an appropriate vector database according to the following decision tree:

你的情况是什么?
│
├─── 已有 PostgreSQL,且数据量 < 500 万
│    └──> 用 pgvector,无缝集成,零额外运维
│
├─── 做 AI/LLM 应用原型,快速验证
│    └──> 用 Chroma,几行代码跑起来
│
├─── 需要生产级部署,性能优先,数据量 500 万 ~ 1 亿
│    └──> 用 Qdrant,Rust 实现,性能强
│
├─── 超大规模(> 1 亿),有 K8s 运维能力
│    └──> 用 Milvus,分布式,功能最全
│
└─── 团队没有运维能力,愿意付费
     └──> 用 Pinecone 云服务,开箱即用

Choosing an embedding model

Choosing a suitable embedding model is the key first step in vector database applications.

RequirementsRecommended Model
Chinese and English text (High Quality)OpenAI text-embedding-3-small
Chinese text (Local Offline)BAAI/bge-large-zh-v1.5
Multilingual universal.paraphrase-multilingual-MiniLM-L12-v2
Image-text multimodal.OpenAI CLIPseries

Performance optimization tips

1. Batch insertion: Insert multiple entries at once to avoid frequent single-entry writes.

Example

# Recommended: Batch Insert
collection.add(documents=docs_list, ids=ids_list)

# Not recommended: looping one by one (re-indexes every time, extremely inefficient)
# for doc, id in zip(docs_list, ids_list):
#     collection.add(documents=[doc], ids=[id])

2. Vector normalization: Before using cosine similarity, normalizing vectors in advance can speed up computation.

Example

import numpy as np

def normalize(v):
    """Perform L2 normalization on vectors so that the norm is 1"""
    return v / np.linalg.norm(v)

3. Set n_results reasonably: Don’t blindly set a very large top_k; 3~10 results are usually enough for RAG scenarios.

4. Make good use of metadata filtering: Combine with where conditions during search to narrow the search scope.

Example

# Only search in the "technical documentation" category to narrow scope and improve accuracy
collection.query(
    query_texts=["Python exception handling"],
    where={"category": "tech_doc"},
    n_results=5
)

5. Rebuild indexes regularly: After data volume grows, rebuild the HNSW index in a timely manner to maintain query performance.

Common pitfalls

The following are the most common problems beginners encounter when using vector databases, along with solutions.

ProblemDescriptionSolution
Embedding model must be consistentThe same model must be used for insertion and queryPin the model version in the configuration file
Dimension mismatchSwitched models but didn’t rebuild the collectionDelete and rebuild the collection when changing models
Text too longMost models have token limits (512~8192)Chunk overly long text first
Inaccurate similarityText is not chunked, so semantics are dilutedSplit documents by paragraph or fixed length
Slow cold startWhen data volume is large, the initial index loading takes timePre-warm in advance, or use persistent indexes

Summary

Let’s review the core knowledge points of this tutorial:

Vector Database Core Knowledge Review Vector / embedding Convert objects into high-dimensional numeric vectors semantically similar means close distance is the foundation of everything Similarity calculation Cosine similarity Euclidean distance Dot product Cosine is preferred for text Indexing algorithms HNSW is the most mainstream IVF suits large scale Flat suits small data Approximation for speed Selection reference Beginner: Chroma Production: Qdrant Existing PG: pgvector 100M+ scale: Milvus Core applications RAG knowledge QA Personalized recommendation Reverse image search Anomaly detection

Vector databases are a crucial part of AI-era infrastructure.

It solves the semantic similarity search problem that traditional databases cannot handle, and is a core component for building RAG systems, recommendation systems, and multimodal search.

Learning path recommendations

  1. Step 1: Understand the concepts of vectors and embeddings, and run the Chroma example from this article.
  2. Step 2: Try building a simple document Q&A system with LangChain + Chroma.
  3. Step 3: Learn indexing algorithms such as HNSW, and understand the trade-off between accuracy and speed.
  4. Step 4: Based on actual project requirements, select an appropriate vector database and deploy it in a production environment.

Reference resources

Other extensions