RAG and Knowledge Retrieval
RAG (Retrieval-Augmented Generation) is currently one of the most mainstream LLM deployment architectures.
The core idea of RAG is:When answering questions, let the LLM first retrieve relevant content from an external knowledge base, then generate a response based on the retrieval results, rather than relying only on knowledge memorized during model training.
This solves two core pain points of LLMs: knowledge cutoff (the model does not know events that occurred after training) and hallucination (the model fabricates answers when uncertain).
Basic Principles of RAG
A complete RAG system consists of two pipelines:Offline indexing pipeline(preprocesses documents and stores them in a vector database) andonline query pipeline(receives user questions, retrieves, and generates).
In the offline stage, raw documents are split into small chunks, converted into vectors via an Embedding model, and stored in a vector database.
In the online stage, the user question is also converted into a vector, the most similar document chunks are found from the database, concatenated into context, and passed to the LLM to generate an answer.
The figure below shows the complete request flow of RAG:
Data Preprocessing and Document Chunking
Prerequisite Challenge: Complex Document Parsing
Before chunking, RAG often facesformat parsingchallenges. Especially for tables, images, and multi-column layouts in PDFs, Word documents, or scanned files, ordinary text extraction can easily cause semantic confusion.
The current mainstream industry approach is to introducedocument parsing engines(such as LlamaParse, Unstructured) or multimodal large models, converting complex images and text into structured Markdown, laying the foundation for subsequent high-quality chunking.
Document Chunking Strategies
Document chunking is the foundation of RAG effectiveness; chunk granularity directly affects retrieval quality. Chunks that are too large introduce noise, while chunks that are too small lose context. Common strategies are as follows:
| Chunking strategy | Applicable scenarios | Advantages | Disadvantages |
|---|---|---|---|
| Fixed-size chunking | General text | Simple to implement, fast | May cut off semantically complete sentences |
| Recursive character chunking | Structured text (Markdown, code) | Prefers splitting along semantic boundaries such as paragraphs and sentences | Slightly complex to implement; requires setting a reasonable list of separators |
| Semantic chunking | Long documents, books | Uses Embedding to compute similarity between adjacent sentences and automatically finds semantic turning points for chunking | High computational cost, slow preprocessing |
| Parent-child document retrieval (Small-to-Big) |
Comprehensive coverage scenarios | Uses "small chunks" for high-precision vector retrieval, and upon a hit returns the corresponding "large chunk" (parent document) to the LLM, balancing retrieval precision and context completeness. | Database design and maintenance costs double |
In practice, it is common to add during chunkingoverlap, i.e., adjacent chunks share several characters to prevent important information from being truncated at boundaries. Typical configuration: chunk size 512 tokens, overlap 50~100 tokens.
Example: Recursive Chunking with LangChain
splitter = RecursiveCharacterTextSplitter(
chunk_size=512, # Maximum number of tokens per chunk
chunk_overlap=50, # Overlap token count between adjacent chunks to prevent information loss at boundaries
separators=["\n\n", "\n", "。", ".", " ", ""] # Prioritize splitting by paragraphs and sentences
)
chunks = splitter.split_text(document_text)
print(f"Split into {len(chunks)} document chunks")
Vector Retrieval
Embedding Models
Embedding models convert text into dense vectors (usually 768- or 1536-dimensional float arrays). Semantically similar texts are closer in vector space, which is the mathematical basis of similarity retrieval.
Comparison of common embedding models:
| Model | Dimensions | Supported languages | Features |
|---|---|---|---|
text-embedding-3-small(OpenAI) |
1536 | Multilingual | High cost-effectiveness, suitable for large-scale indexing |
text-embedding-3-large(OpenAI) |
3072 | Multilingual | Highest accuracy, higher cost |
BAAI/bge-m3 |
1024 | Chinese and English | Open source, excellent Chinese performance, supports multiple languages |
sentence-transformers/all-MiniLM-L6-v2 |
384 | English | Small size, fast speed, suitable for extremely lightweight local deployment |
Similarity Computation and ANN Algorithms
The core of retrieval is measuring distance. The most commonly used isCosine Similarity, which computes the cosine of the angle between two vectors, with a range of [-1, 1]; the closer to 1, the more similar. Additionally, there are Dot Product and Euclidean Distance (L2 Distance).
To achieve millisecond-level retrieval among millions of vectors, databases typically useApproximate Nearest Neighbor (ANN) algorithms(such asHNSW, IVF). HNSW is currently the most mainstream algorithm. It builds a multi-layer skip graph network, sacrificing very little accuracy in exchange for an order-of-magnitude improvement in search speed.
Advanced RAG (Advanced Architecture)
The basic architecture (Naive RAG) often faces issues such as inaccurate retrieval and "context flooding" caused by redundant information. Advanced RAG addresses these throughpre-retrieval optimization → retrieval fusion → post-retrieval optimizationusing a three-stage architecture to solve them.
1. Pre-retrieval: Query Optimization
Users' original questions are often not precise:
- Query Rewriting: Use LLM to rewrite colloquial queries into standardized retrieval terms.
- HyDE(Hypothetical Document Embedding): Let the LLM first "blind-guess" a hypothetical answer. Since the generated answer usually contains more industry terminology than the original question, using the vector of this hypothetical answer for retrieval often recalls higher-quality documents.
2. Hybrid Search
willVector Retrieval(Understands semantics, high fault tolerance) andKeyword RetrievalThe results are fused by weight. This is especially important when encountering proper nouns, product models, and code snippets, because traditional vector retrieval is prone to "failing" on specific proper nouns.
3. Post-retrieval Optimization: Reranking
This is acoarse ranking → fine rankingtwo-stage design. Although vector retrieval is fast, its scoring is not precise enough. Reranking introducesCross-Encoder model(such as `bge-reranker`), which inputs "question" and "document" pairs into the model for joint inference and scoring. It is computationally heavy and only responsible for selecting the Top-20 down to Top-5.
Example: Reranking Pipeline Pseudocode
reranker = CrossEncoder("BAAI/bge-reranker-v2-m3")
# 1. Coarse ranking: vector retrieval rapidly recalls Top-50
candidates = vector_store.similarity_search(query, k=50)
# 2. Fine ranking: construct [question, document] pairs for precise scoring
pairs = [[query, doc.page_content] for doc in candidates]
scores = reranker.predict(pairs)
# 3. Filter the final Top-5 to pass into the LLM
ranked_docs = sorted(zip(scores, candidates), reverse=True)
final_docs = [doc for _, doc in ranked_docs[:5]]
4. Self-RAG and CRAG (Corrective RAG)
A self-reflection mechanism is added. For example, CRAG (Corrective RAG), after obtaining retrieval results, first has the LLM act as a "judge" to score them. If the local knowledge base has no matching document or the quality is extremely low, the system automatically triggers Web Search (such as Google API) as a supplement, greatly reducing hallucinations.
GraphRAG: Knowledge Graph + Retrieval Integration
Traditional RAG treats the knowledge base as independent text fragments and cannot answer questions such as "find all companies founded by the current CEO and with a market value exceeding 100 billion" that requirecross-document, multi-hop reasoningcomplex problems.GraphRAGIntroducing a Knowledge Graph to explicitly model entities and relationships.
GraphRAG Core Steps
- Knowledge Construction: In the offline phase, use LLM to extract triples (subject, relationship, object) from documents and write them into graph databases such as Neo4j.
- Dual-Path Retrieval: For entities in the query, not only is traditional vector retrieval performed, but graph traversal is also triggered in the knowledge graph to extract multi-hop relationship chains.
- Graph-Text Fusion Generation: Assemble the "chunks" retrieved by vector retrieval and the "path structures" retrieved by graph retrieval into the Prompt, giving the LLM both a global view and specific details.
GraphRAG Content Reference:https://www.example.com/ai-agent/graphrag-usage.html
Technology and Database Selection Recommendations
| Database/Tool Selection | Type | Recommended Implementation Scenarios |
|---|---|---|
| Pinecone / Zilliz Cloud | Fully Managed Cloud Service | Out-of-the-box, no need to maintain infrastructure. Combined with Cohere Rerank + GPT-4o, it is the fastest commercial solution. |
| Qdrant | Open Source + Managed | Written in Rust, excellent memory management, extremely high performance. Suitable for enterprise-level private deployment. |
| Weaviate / Elasticsearch | Open Source + Managed | Comes with an extremely mature BM25 + vector hybrid search (Hybrid Search), and is the top choice for scenarios with many proper nouns. |
| Milvus | Open Source Distributed | Suitable for ultra-large-scale enterprise retrieval platforms at the billion to ten-billion level. |
| Chroma / FAISS | Local Library/Embedded | Extremely lightweight, no need to deploy a standalone service. Very suitable for local development and personal knowledge base project validation. |
RAG Evaluation Metrics (RAGAS Framework)
RAG system evaluation cannot rely on intuition alone; it mainly usesRAGASa framework for automated quantitative testing from the two dimensions of "retrieval" and "generation":
- Context Recall: What proportion of the information in the standard answer can be retrieved.
- Context Precision: What proportion of the retrieved documents are truly relevant.
- Faithfulness (Faithfulness/Hallucination Metric): Whether the generated answers are all supported by the retrieved documents.
- Answer Relevance: Whether the generated answer truly addresses the user's question, avoiding irrelevant responses.