LangChain RAG

RAG (Retrieval-Augmented Generation) allows AI to answer questions based on your private documents without fine-tuning the model. You only need to vectorize and store the documents, and the Agent can retrieve relevant content to answer.


What is RAG

Ordinary large models can only answer content that exists in their training data. If your documents are private (company internal documents, personal notes), the model "doesn't know" about them. RAG solves this problem:

  • Offline stage: Split documents into small chunks → Use an Embedding model to convert them into vectors → Store them in a vector database
  • Online stage: User asks a question → Convert the question into a vector → Search for the most similar content in the vector database → Send the retrieved content to the model as context → The model answers based on the retrieved content

RAG Workflow

┌──────────────────────────────────────────────────────┐
│                    离线阶段(索引)                     │
│                                                      │
│  文档 → DocumentLoader → TextSplitter → Embedding     │
│                                           ↓          │
│                                      VectorStore     │
└──────────────────────────────────────────────────────┘

┌──────────────────────────────────────────────────────┐
│                    在线阶段(检索)                     │
│                                                      │
│  用户提问 → Embedding → 相似度搜索 → 检索结果          │
│                                          ↓           │
│                      检索结果 + 用户问题 → 模型 → 回答  │
└──────────────────────────────────────────────────────┘


Environment Preparation

Install RAG-related dependencies:

$ pip install langchain-deepseek langchain-chroma chromadb
PackagePurpose
langchain-deepseekProvides the OpenAI Embedding model
langchain-chromaLangChain integration for the Chroma vector database
chromadbChroma vector database (lightweight, suitable for beginners)

Embedding Model Initialization

Example

from dotenv import load_dotenv
load_dotenv()

from langchain_openai import OpenAIEmbeddings

# OpenAI's text embedding model
# Convert text into a vector (a set of floating-point numbers)
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

# Test: convert a piece of text into a vector
text = "EXAMPLE is a programming learning platform"
vector = embeddings.embed_query(text)

print(f"Text: {text}")
print(f"Vector dimension: {len(vector)}")   # text-embedding-3-small is 1536-dimensional
print(f"First 5 values of the vector: {vector[:5]}")

Run result:

Text: Example 是一个编程学习平台
Vector dimension: 1536
First 5 vector values: [0.0123, -0.0045, 0.0234, -0.0012, 0.0089]

Alternatively, if you don't have an OpenAI key, you can use Alibaba Bailian's Embedding service..envYou need to add the Alibaba Bailian key to the configuration file:

DASHSCOPE_API_KEY="sk-xxx"

Code after configuration:

Example

import os
from dotenv import load_dotenv
load_dotenv()

from langchain_openai import OpenAIEmbeddings

# OpenAI's text embedding model
# Convert text into a vector (a set of floating-point numbers)

# Use Alibaba Cloud Bailian (DashScope) Qwen Embedding service
# Bailian's Embedding API is compatible with the OpenAI API specification, so directly use langchain-openai
# OpenAIEmbeddings, just point base_url to Bailian's compatible endpoint,
# no need to install langchain-community / dashscope (this package is no longer maintained).
# text-embedding-v4 is currently the recommended general-purpose vector model, outputting 1024-dimensional vectors by default.
#
# Two key parameters are essential:
# - check_embedding_ctx_length=False: OpenAIEmbeddings uses tiktoken by default
# to pre-encode text into token id arrays before sending (this is the format recognized by the official OpenAI API),
# but Bailian's compatible interface only accepts raw strings. If you don't disable this option, it will report
# a "contents is neither str nor list of str" error.
# - chunk_size=10: Bailian's Embedding interface accepts at most 10 pieces of text per request.
# OpenAIEmbeddings packs 1000 pieces at a time by default. If the knowledge base is slightly larger, it will exceed the limit and report an error.
embeddings = OpenAIEmbeddings(
    model="text-embedding-v4",
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
    check_embedding_ctx_length=False,
    chunk_size=10,
)
# Test: convert a piece of text into a vector
text = "EXAMPLE is a programming learning platform"
vector = embeddings.embed_query(text)

print(f"Text: {text}")
print(f"Vector dimension: {len(vector)}")   # text-embedding-3-small is 1536-dimensional
print(f"First 5 values of the vector: {vector[:5]}")

Create Vector Store

Example

from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma

# Initialize the Embedding model
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

# Create a Chroma vector store (data is saved in a local directory)
vector_store = Chroma(
    collection_name="example_docs",
    embedding_function=embeddings,
    persist_directory="./chroma_db",  # Persistence directory
)

# Add documents (the simplest form: a list of texts)
texts = [
    "EXAMPLE is a free programming learning website that provides tutorials on HTML, CSS, JavaScript, Python, etc.",
    "The Python3 basic tutorial has 30 chapters, suitable for beginners with zero foundation, covering environment setup, basic syntax, object-oriented programming, and more.",
    "The HTML basic tutorial has 25 chapters, covering basic knowledge such as HTML tags, forms, and multimedia.",
]

# add_texts automatically converts text into vectors and stores them
vector_store.add_texts(texts)

print(f"Added {len(texts)} documents to the vector store")

Semantic Search

Example

# Semantic search — relies not on keyword matching but on semantic similarity
results = vector_store.similarity_search(
    "I want to learn Python. What tutorials do you recommend?",
    k=2,  # Return the 2 most similar results
)

print("Search results:")
for i, doc in enumerate(results):
    print(f"\n"Result {i+1}:")
    print(f" Content: {doc.page_content}")
    print(f" Metadata: {doc.metadata}")

Run result:

搜索结果:

Result 1:
  Content: Python3 基础教程共 30 章,适合零基础入门...
  Metadata: {}

Result 2:
  Content: Example是一个免费的编程学习网站...
  Metadata: {}

Note that the first search result is more relevant than the second — although the first contains the keyword "Python", it is ranked bysemantic similarityrather than keyword matching. This is the advantage of vector retrieval.


Create a Retriever

Retriever is a standardized interface for Vector Store:

Example

# Create a retriever from vector_store
retriever = vector_store.as_retriever(
    search_type="similarity",  # Similarity search
    search_kwargs={"k": 3},    # Return the top 3 results
)

# Use the retriever
docs = retriever.invoke("Python Learning Path")
for doc in docs:
    print(f"- {doc.page_content[:60]}...")
Other Extensions