Python Implementation of RAG and Knowledge Retrieval
This chapter introduces Retrieval-Augmented Generation (RAG) technology.
RAG is one of the core technologies for building knowledge-intensive agents.
By combining vector retrieval with language generation models, RAG gives agents the ability to access and utilize external knowledge, rather than relying solely on content "remembered" in model parameters.
RAG Fundamentals
The core idea of RAG is to combine two stages, "information retrieval" and "language generation", to form a complete question-answering pipeline.
When a user asks a question, the system first retrieves content relevant to the question from the knowledge base.
Then, the system provides the retrieved content as context to the generation model, helping it produce more accurate and well-supported answers.
Why RAG is Needed
Language models have a knowledge cutoff date and cannot perceive events that occurred after training was completed.
Model parameters are finite, and it is impossible to memorize all the world's knowledge in them.
If the model relies entirely on "recalling" knowledge from parameters, when it encounters content it cannot remember precisely or has never learned, it often fabricates answers with a straight face. This is what is commonly called "hallucination".
RAG's approach is to have the model consult real data before answering, turning "answering from memory" into "open-book answering". The knowledge source can also be updated at any time without retraining the model.
Core Architecture
A complete RAG system usually consists of three core components:
Retriever: Responsible for finding the most relevant information snippets from the knowledge base.
Vector Store: Stores vector representations of document content, supporting large-scale, low-latency similarity search.
Generator: Generates the final natural language answer based on the retrieved content and the user's original question.
The relationship among the three can be understood as: the retriever is responsible for "finding materials", the vector store is responsible for "storing materials and quickly finding similar materials", and the generator is responsible for "organizing the answer while looking at the materials".
Workflow
Step 1:Indexing phase. Split the original documents into several text chunks, vectorize them one by one, and store them in the vector database. This step is usually done offline and in advance.
Step 2:Retrieval phase. When a user asks a query, the query is also vectorized, and a similarity search is performed in the vector database to find the most relevant text chunks.
Step 3:Augmentation phase. Concatenate the retrieved relevant text chunks with the user's original question to construct a more complete prompt.
Step 4:Generation phase. The generation model produces the final answer based on this augmented context.
Among these four steps, only the first step is offline preprocessing; the other three happen in real time when the user asks. Therefore, retrieval speed and retrieval quality directly determine the response speed and answer quality of the entire system.
Code Implementation
Basic RAG Implementation
"""
Simplified RAG System Implementation
Contains three core components: embedder, vector store, generator
"""
def __init__(self, embedder, vector_store, generator):
# Embedder: responsible for converting text into vectors
self.embedder = embedder
# Vector database: responsible for storing vectors and supporting similarity search
self.vector_store = vector_store
# Generator: generates answers based on context
self.generator = generator
def add_documents(self, documents):
"""
Add documents to knowledge base
:param documents: original document list (each element is a complete piece of text)
"""
# Step 1: Split long documents into smaller text chunks for precise retrieval
chunks = self.chunk_documents(documents)
# Step 2: Convert all text chunks into vectors
embeddings = self.embedder.embed(chunks)
# Step 3: Store the vectors together with the corresponding original text in the vector database
self.vector_store.add(embeddings, chunks)
def chunk_documents(self, documents, chunk_size=500, overlap=50):
"""
Document splitting: split long documents into text chunks of fixed length with some overlap.
:param documents: raw document list
:param chunk_size: number of characters per text chunk
:param overlap: number of overlapping characters between adjacent text chunks (to avoid semantics being cut off at the split point)
:return: list of text chunks after splitting
"""
chunks = []
for doc in documents:
# Sliding split with a stride of (chunk_size - overlap), retaining a certain overlapping region
for i in range(0, len(doc), chunk_size - overlap):
chunk = doc[i:i + chunk_size]
chunks.append(chunk)
return chunks
def query(self, question, top_k=5):
"""
Process user query and return the generated answer
:param question: question asked by the user
:param top_k: number of relevant text chunks returned during retrieval
:return: final answer given by the generation model
"""
# 1. Convert the question into a vector so it can be compared with vectors in the knowledge base
question_embedding = self.embedder.embed([question])[0]
# 2. Retrieve the top_k most similar text chunks from the vector database
results = self.vector_store.search(question_embedding, top_k)
# 3. Concatenate the retrieved text chunks into context
context = "\n".join([doc for doc, _ in results])
# 4. Organize the context and the original question into a prompt
prompt = f"Answer the question based on the following content:\n{context}\n\nQuestion: {question}"
# 5. Hand it to the generation model to produce the final answer
return self.generator.generate(prompt)
This basic implementation can already run the complete RAG pipeline, but it is often not enough in real projects: the chunking method is too crude, and retrieval ranking relies only on a single vector similarity calculation, resulting in limited accuracy. Below are some common optimization ideas.
Advanced RAG
Common problems in basic RAG include: poor quality of retrieved content, insufficient ranking accuracy, and irrelevant information mixed into the context.
Advanced RAG further optimizes the basic pipeline by introducing techniques such as reranking and hybrid retrieval, thereby improving the accuracy of the final answer.
Reranking
Vector retrieval in basic RAG is fast, but ranking accuracy is limited, and it can easily rank less semantically relevant content higher.
The reranking approach is: first use vector retrieval to quickly filter out a set of candidate texts (e.g., top-20), then use a more computationally expensive but more accurate cross-encoder (Cross-Encoder) to score each candidate one by one, re-rank them, and extract the truly most relevant few (e.g., top-5).
This two-stage approach of "rough filtering first, then precise ranking" balances retrieval speed and retrieval accuracy.
RAG with Reranking
"""
Advanced RAG system
Add a reranking step on top of the basic pipeline
"""
def __init__(self, embedder, vector_store, reranker, generator):
# Embedder
self.embedder = embedder
# Vector database
self.vector_store = vector_store
# Reranking model (cross-encoder) for fine-grained ranking
self.reranker = reranker
# Generator
self.generator = generator
def query(self, question, initial_k=20, final_k=5):
"""
Process query, including two-stage retrieval of "rough filtering + precise ranking"
: :param question: user question
: :param initial_k: The number of candidates returned by the first stage of vector retrieval.
:param final_k: The number of items ultimately retained after re-ranking.
"""
# ==================== Phase 1: Initial Retrieval (Coarse Screening) ====================
Convert the problem into a vector.
question_embedding = self.embedder.embed([question])[0]
# Quickly recall a batch of candidate texts using vector similarity (fast, average accuracy)
initial_results = self.vector_store.search(
question_embedding,
initial_k
)
# ==================== Phase 2: Re-ranking (Fine Ranking) ====================
Construct (question, candidate document) pairs and feed them to the cross-encoder for scoring one by one.
doc_pairs = [(question, doc) for doc, _ in initial_results]
Cross-encoders encode the question and document simultaneously, allowing for more accurate judgment of their relevance.
reranked = self.reranker.rerank(doc_pairs)
# Take the truly most relevant final_k items after re-ranking.
final_context = "\n".join([doc for doc, _ in reranked[:final_k]])
# ==================== Stage 3: Generation =====================
prompt = fContext: {final_context}\n\nQuestion: {question}
return self.generator.generate(prompt)
Hybrid Retrieval
Hybrid Retrieval combines the advantages of dense retrieval and sparse retrieval.
Dense Retrieval is based on vector similarity and excels at capturing semantic relevance, matching user queries and documents even when they do not use exactly the same words.
Sparse retrieval (such as BM25) is based on keyword frequency statistics, excels at exact literal matching, and is often more reliable than vector retrieval for scenarios such as proper nouns, numbers, and code snippets.
Both have blind spots: vector retrieval may "understand the meaning but fail to find the exact keywords," while keyword retrieval may "match the words but fail to understand the semantics." Hybrid retrieval fuses the results of the two retrieval paths, complementing each other's strengths and weaknesses.
Hybrid Retrieval RAG
"""
Hybrid retrieval RAG
Combining dense retrieval (vector) and sparse retrieval (keyword)
"""
def __init__(self, dense_retriever, sparse_retriever, fusion_fn):
# Dense Retriever: Based on vector similarity, excels at semantic matching.
self.dense_retriever = dense_retriever
# Sparse Retriever: Based on keyword statistics (e.g., BM25), excels at exact matching.
self.sparse_retriever = sparse_retriever
# Fusion function: used to merge two retrieval results, e.g., RRF (see implementation below)
self.fusion_fn = fusion_fn
def query(self, question, top_k):
"""
Perform hybrid retrieval and return the fused ranked results.
: :param question: user question
:param top_k: The final number of results to return.
"""
# ==================== Dense Retrieval =====================
# Retrieve more candidates (top_k * 2) to leave room for filtering during fusion
dense_results = self.dense_retriever.search(question, top_k * 2)
# ==================== Sparse Retrieval ====================
# Keyword-based retrieval, also retrieve more candidates
sparse_results = self.sparse_retriever.search(question, top_k * 2)
# ==================== Result Fusion ====================
# Use strategies like Reciprocal Rank Fusion (RRF) to fuse the ranking results of both channels
fused = self.fusion_fn(dense_results, sparse_results)
# Take the top_k items with the highest fused ranking
return fused[:top_k]
def reciprocal_rank_fusion(results_a, results_b, k=60):
"""
Reciprocal Rank Fusion (RRF) algorithm
Used to fuse the ranking results from two (or multiple) retrieval channels into a unified ranking
Core idea: only look at the "rank" of each document in each channel's results, not the specific score,
This avoids the problem that vector similarity scores and BM25 scores have different scales and cannot be directly compared.
:param results_a: Retrieval results A, formatted as [(doc, score), ...]
:param results_b: Retrieval results B, formatted as [(doc, score), ...]
:param k: Smoothing constant, used to reduce the weight difference for lower-ranked documents, usually set to 60
:return: List of [(doc, score), ...] sorted by fused score from high to low
"""
scores = {}
# Process results A: the higher the rank, the higher the score contributed
for rank, (doc, _) in enumerate(results_a):
scores[doc] = scores.get(doc, 0) + 1 / (k + rank + 1)
# Process results B, similarly accumulate scores by rank
for rank, (doc, _) in enumerate(results_b):
scores[doc] = scores.get(doc, 0) + 1 / (k + rank + 1)
# Sort by the accumulated fused score from high to low
sorted_docs = sorted(scores.items(), key=lambda x: -x[1])
return sorted_docs
Other Optimization Strategies
Query expansion: Expand the original query by adding synonyms or related terms to broaden the recall range and improve recall rate.
Query rewriting: Use models to understand the user's true intent, rewriting colloquial or ambiguous questions into formulations more suitable for retrieval.
Context compression: Before passing retrieval results to the generation model, remove redundant and irrelevant parts, keeping only key information. This reduces cost and also lowers the risk of the generation model being disturbed by irrelevant content.
Vector Database
A vector database is the infrastructure of the RAG system, responsible for storing massive high-dimensional vectors and supporting large-scale, low-latency similarity retrieval.
Core Concepts
Embedding: The process of converting text (or images, audio, etc.) into a set of dense numerical vectors, where semantically similar content is closer in vector space.
Vector dimensionThe length of the embedding vector. The higher the dimension, the richer the semantic information it can theoretically express, but storage and computational costs will also increase accordingly.
Similarity measure.: A specific calculation method used to measure the "similarity" or "distance" between two vectors. Different metrics are suitable for different scenarios.
Common Similarity Metrics
| measurement method | formula | Features | Applicable scenarios |
|---|---|---|---|
| Cosine similarity. | cos(A,B) = A·B/(|A||B|) | Only measures directional similarity, independent of vector length, with a value range of [-1,1] | Document/text similarity |
| Euclidean distance | d(A,B) = sqrt(sum((Ai-Bi)^2)) | Measures the absolute spatial distance between vectors, with a value range of [0,+∞) | Image feature matching |
| dot product | A·B = sum(Ai*Bi) | Affected by both direction and length, with high computational efficiency | recommendation system |
Comparison of Mainstream Vector Databases
| database | Type | Features | Applicable scenarios |
|---|---|---|---|
| Pinecone | cloud service | Fully managed service, easy to get started, auto-scaling, no manual operations required | Production environments that need to go live quickly |
| Weaviate | Open source | Natively supports hybrid retrieval (vector + keyword), highly modular | Scenarios requiring flexible customization of retrieval logic |
| Milvus | Open source | High availability, easy horizontal scaling, can support vector scales of billions and above | Ultra-large-scale vector search |
| Chroma | Open source | Lightweight, simple API, can be up and running with just a few lines of code | Prototype development, teaching, and local testing |
| Qdrant | Open source | Built on Rust, excellent performance, supports rich filtering conditions | Scenarios requiring high throughput and complex filtering |
Code Examples
Using Chroma Vector Database
from chroma_client import ChromaClient
# Create a client (can connect to a local instance or a remote service)
client = ChromaClient()
# Create or get a collection, equivalent to a 'table' in a relational database
collection = client.get_or_create_collection("knowledge_base")
# Add documents: ID, vector, original text, and metadata must correspond one-to-one
collection.add(
ids=["doc1", "doc2", "doc3"], # Unique identifier for each document
embeddings=[
[0.1, 0.2, 0.3], # Vector corresponding to document 1 (the actual dimension is usually much higher than 3 dimensions)
[0.4, 0.5, 0.6], # Vector corresponding to document 2
[0.7, 0.8, 0.9], # Vector corresponding to document 3
],
documents=[
Artificial intelligence is a branch of computer science.,
"Machine learning is a subfield of artificial intelligence",
"Deep learning is a subfield of machine learning"
],
metadatas=[
{"source": “textbook”}, # Source information of Document 1, can be used for subsequent filtering
{"source": Paper},
{"source": Paper}
]
)
# Retrieve: pass in query vector, return the most similar n_results results
results = collection.query(
query_embeddings=[[0.15, 0.25, 0.35]], # Vector corresponding to the query
n_results=2 # Return top-2 most similar results
)
# Output retrieval results
print(results["documents"]) # Matched original text content
print(results["distances"]) # Corresponding distance/similarity score; smaller values usually indicate higher similarity
When choosing a vector database, you need to comprehensively consider data scale, query latency requirements, deployment method (cloud-hosted or self-hosted), and the team's operational capabilities. For small-scale prototype development, prefer Chroma; for a production environment that is out-of-the-box, choose Pinecone; for ultra-large-scale scenarios requiring autonomy and control, Milvus or Qdrant are more suitable.
GraphRAG
GraphRAG is an advanced solution that combines Knowledge Graph with RAG.
It enhances the system's retrieval and reasoning capabilities by constructing a graph structure composed of entities and relations.
Why GraphRAG is Needed
Traditional RAG is based on similarity retrieval over text chunks, making it difficult to handle problems that require "multi-hop reasoning" (i.e., answers that require connecting multiple pieces of information).
Traditional RAG also has difficulty explicitly capturing semantic relationships between entities, such as structured information like "who is whose superior" or "which company acquired which company."
GraphRAG explicitly organizes these entities and relations into a graph structure, allowing the system to reason along relation paths instead of just "finding similar text."
Core Advantages
Can capture semantic relationships between entities, forming a structured knowledge network rather than scattered text fragments.
Supports path-based reasoning and can better answer relational questions.
Can handle complex problems requiring multi-hop reasoning, such as "who is someone's advisor's advisor," which require traversing multiple relation nodes in sequence to derive the answer.
Code Implementation
GraphRAG Implementation
"""
GraphRAG system
Combines knowledge graph retrieval and vector retrieval, balancing relational reasoning and semantic similarity.
"""
def __init__(self, kg_builder, embedder, vector_store, generator):
# Knowledge graph builder: responsible for entity/relation extraction and graph queries
self.kg_builder = kg_builder
# Embedder: used to convert documents and questions into vectors
self.embedder = embedder
# Vector database: serves as a supplementary channel for knowledge graph retrieval
self.vector_store = vector_store
# Generator: generates the final answer based on the fused context
self.generator = generator
def index_documents(self, documents):
"""
Index documents into both the knowledge graph and the vector database
:param documents: list of documents
"""
for doc in documents:
# Extract entities from documents
# For example, extract from "OpenAI was founded in 2015"
# Entity: OpenAI (organization), 2015 (time)
entities = self.kg_builder.extract_entities(doc)
# Extract relationships between entities from documents
# For example, extract a relation triplet: (OpenAI, founded in, 2015)
relations = self.kg_builder.extract_relations(doc)
# Write entities and relationships into the knowledge graph
self.kg_builder.add_entities(entities)
self.kg_builder.add_relations(relations)
# Simultaneously vectorize the original text and store it in a vector database to supplement semantic retrieval
embedding = self.embedder.embed([doc])
self.vector_store.add(embedding, [doc])
def query(self, question):
"""
Handle queries: combine knowledge graph reasoning and vector retrieval
"""
# ==================== Knowledge Graph Retrieval ====================
# First identify the relevant entities mentioned in the question
# For example, "Who founded OpenAI?" -> identify the entity "OpenAI"
relevant_entities = self.kg_builder.query_entities(question)
# Based on these entities, extract the subgraph directly related to them from the graph
# The subgraph contains relevant entities and their adjacent nodes and connecting edges
subgraph = self.kg_builder.get_subgraph(relevant_entities)
# ==================== Vector Retrieval ====================
# As a supplementary channel, retrieve semantically similar text content
vector_results = self.vector_store.search(question, top_k=5)
# ==================== Result Fusion ====================
# Fuse the graph reasoning results and vector retrieval results into a unified context
combined_context = self.fuse_results(subgraph, vector_results)
# ==================== Generation ====================
return self.generator.generate(combined_context, question)
def fuse_results(self, subgraph, vector_results):
"""
Fuse the knowledge graph subgraph and vector retrieval results into context text readable by a generative model
"""
# Convert the subgraph into natural language descriptions (e.g., "OpenAI -founded in-> 2015")
kg_text = subgraph.to_text()
# Concatenate the original text segments retrieved by vector retrieval
vector_text = "\n".join([doc for doc, _ in vector_results])
return f"{kg_text}\n{vector_text}"
class KnowledgeGraphBuilder:
"""Knowledge Graph Builder (illustrative implementation; in real projects, dedicated NER/relation extraction models are usually called)"""
def extract_entities(self, text):
"""
# Extract named entities from text, usually done with NER (Named Entity Recognition) models
# Common entity types include PERSON (person), ORG (organization), TIME (time), etc.
"""
entities = []
return entities
def extract_relations(self, text):
"""
# Extract relationships between entities from text, usually done with relation extraction models
# Relationships are generally represented as triplets: (head entity, relation type, tail entity)
"""
relations = []
return relations
def query_entities(self, question):
"""Based on the user question, identify the entities involved in the question for subsequent subgraph retrieval"""
return []
Application Scenarios
GraphRAG is especially suitable for the following scenarios:
Question answering that requires understanding entity relationships, e.g., "Who is the CEO of this company?"
Questions that require multi-hop reasoning to arrive at an answer, e.g., "Who is the advisor of Zhang San's advisor?"
Scenarios that require mining implicit relationships in the knowledge base that are not explicitly written.
It should be noted that building and maintaining a knowledge graph itself incurs certain engineering costs (entity/relation extraction accuracy, graph update and maintenance, etc.), so it is more suitable for scenarios with clear relational reasoning requirements, and is not recommended as the default starting point for all RAG projects.
Solution Comparison and Recommendations
| Solution | Core Approach | Advantages | Limitations | Applicable Scenarios |
|---|---|---|---|---|
| Basic RAG | Vector retrieval + direct generation | Simple to implement, low cost, quick to get started | Limited retrieval precision, prone to recalling content that is not relevant enough | Simple Q&A, rapid prototype validation |
| Advanced RAG (reranking) | Vector coarse filtering + cross-encoder fine ranking | Significantly improves the precision of retrieval ranking | Adds an extra model inference pass, increasing latency and cost | Scenarios with high requirements for answer accuracy |
| Hybrid Retrieval | Fusion of vector retrieval and keyword retrieval | Balances semantic matching and precise keyword matching | Requires maintaining two retrieval indexes simultaneously | Knowledge bases with many proper nouns, code, and identifiers |
| GraphRAG | Knowledge graph + vector retrieval fusion | Supports relational reasoning and multi-hop Q&A | High graph construction and maintenance costs | Scenarios with clear relational and multi-hop reasoning requirements |
Chapter Summary
This chapter introduces the core technologies of RAG and knowledge retrieval.
RAG Basic Principles: By combining retrieval with generation, it enables agents to answer questions using external knowledge rather than relying entirely on memory in model parameters.
Advanced RAG: Through techniques such as reranking and hybrid retrieval, it further improves retrieval quality and answer accuracy on top of the basic pipeline.
Vector Database: It serves as the infrastructure of RAG systems, responsible for efficiently storing and retrieving vector representations of documents; the choice of solution should be made based on data scale and deployment method.
GraphRAG: Combined with knowledge graphs, it enhances relational reasoning and multi-hop Q&A capabilities, but also brings additional construction and maintenance costs.
The key to choosing an appropriate RAG solution is to first understand the real needs of the business scenario: for simple Q&A scenarios, basic RAG is often sufficient; when higher retrieval precision is required, reranking can be introduced; when the knowledge base contains many proper nouns/keywords, hybrid retrieval is suitable; when a large amount of relational reasoning and multi-hop Q&A is involved, consider introducing GraphRAG. The more complex the technology, the higher the engineering and maintenance costs, so choosing based on actual needs is more important than blindly pursuing "more advanced solutions".
Other Extensions