LangChain Personal Knowledge Base Q&A System
This section builds a personal knowledge base system that can load Markdown files and PDF documents and answer questions based on their content.
System Design
- Document Loading: Supports multiple formats including Markdown, TXT, and PDF
- Vector Retrieval: Chroma persistent storage, supports incremental updates
- Source Citations: Answers include the source document and snippet location
- Streaming Output: Displays the answer token by token
Complete Code
Before running, you need to configure the following.envin the file:DEEPSEEK_API_KEY(Chat model) andDASHSCOPE_API_KEY(Alibaba Cloud Bailian, used for knowledge base text vectorization). For specific application procedures and common troubleshooting, seethe "LangChain Intelligent Customer Service Robot"section.
Example
# pip install langchain langchain-deepseek langchain-openai langchain-community langchain-chroma chromadb pypdf
from dotenv import load_dotenv
load_dotenv()
import os
from pathlib import Path
from langchain.tools import tool
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage
from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import TextLoader, PyPDFLoader
class KnowledgeBase:
"""Personal knowledge base manager"""
def __init__(self, persist_dir: str = "./my_knowledge_db"):
self.persist_dir = persist_dir
# Alibaba Cloud Bailian's Embedding interface is compatible with the OpenAI spec; just call it with langchain-openai
# 's OpenAIEmbeddings — no need to install the DashScopeEmbeddings in langchain-community,
# which has stopped being maintained.
# check_embedding_ctx_length=False: turn off tiktoken pre-tokenization and send the original text directly
# (Bailian's interface does not accept token id arrays).
# chunk_size=10: Bailian's Embedding interface accepts at most 10 texts per request.
self.embeddings = OpenAIEmbeddings(
model="text-embedding-v4",
api_key=os.getenv("DASHSCOPE_API_KEY"),
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
check_embedding_ctx_length=False,
chunk_size=10,
)
self.text_splitter = RecursiveCharacterTextSplitter(
chunk_size=500, chunk_overlap=50,
separators=["\n\n", "\n", "。", "!", "?", ". ", "! ", "? ", " "],
)
self.vector_store = None
self._load_or_create()
def _load_or_create(self):
"""Load an existing vector store or create a new one (Chroma works with the same set of parameters;
whether the directory already exists only affects the prompt message, not the creation logic)"""
is_existing = os.path.exists(self.persist_dir) and os.listdir(self.persist_dir)
self.vector_store = Chroma(
embedding_function=self.embeddings,
persist_directory=self.persist_dir,
)
if is_existing:
print(f"Loaded vector store: {self.vector_store._collection.count()} document chunks")
else:
print("Created a new vector store")
def add_file(self, file_path: str) -> int:
"""Add a file to the knowledge base and return the number of document chunks added.
Automatically choose the appropriate Loader based on the file extension:
- For .pdf use PyPDFLoader (depends on the pypdf library to parse PDFs, returns multiple Documents by page)
- For the rest (.md, .txt and other plain text files) use TextLoader to read the raw text
"""
suffix = Path(file_path).suffix.lower()
if suffix == ".pdf":
loader = PyPDFLoader(file_path)
else:
loader = TextLoader(file_path, encoding="utf-8")
docs = loader.load()
# Uniformly add the file name as the source; each Document loaded by PyPDFLoader also comes with
# a page field (page number), which can also be used to locate the snippet position during retrieval
for doc in docs:
doc.metadata["source"] = Path(file_path).name
chunks = self.text_splitter.split_documents(docs)
self.vector_store.add_documents(chunks)
print(f"Added {Path(file_path).name}: {len(chunks)} document chunks")
return len(chunks)
def add_text(self, text: str, source: str = "Manually add") -> int:
"""Directly add text to the knowledge base"""
chunks = self.text_splitter.create_documents(
[text], metadatas=[{"source": source}]
)
self.vector_store.add_documents(chunks)
return len(chunks)
def search(self, query: str, k: int = 3) -> list:
"""Search the knowledge base"""
return self.vector_store.similarity_search(query, k=k)
def get_retriever(self):
"""Get the retriever"""
return self.vector_store.as_retriever(search_kwargs={"k": 3})
Currently,
TextLoader、PyPDFLoaderthis type of local file loader doesn't yet have a standalone maintained package; it can still only belangchain_community.document_loadersimported from langchain_community. At runtime, you may see a DeprecationWarning for the whole langchain-community package — this warning can be ignored for now; just migrate after the official standalone document loader package is released. This is different from the embedding case above, where a ready-made replacement already existed but wasn't adopted.
Continuing with the Agent section:
Example
kb = KnowledgeBase("./my_knowledge_db")
# Add some sample knowledge
kb.add_text(
"EXAMPLE's Python3 basic tutorial contains the following chapters:"
"1. Python Introduction and Environment Setup 2. Basic Data Types 3. Operators and Expressions "
"4. Conditional Statements if-else 5. Loops for/while 6. Function Definition and Calling "
"7. Modules and Packages 8. File Operations 9. Exception Handling 10. Object-Oriented Programming",
source="Python3 Tutorial Outline"
)
kb.add_text(
"To become a good Python developer, it is recommended to study along the following route:"
"First, master Python basic syntax (1-2 weeks);"
"Second, learn the basics of data structures and algorithms (2-3 weeks);"
"Third, choose a direction for in-depth study (Web Development/Data Analysis/AI);"
"Fourth, work on 2-3 hands-on projects to consolidate your knowledge.",
source="Python Learning Path"
)
kb.add_text(
"EXAMPLE's online programming environment supports multiple languages including Python, JavaScript, Java, and C++."
"Users don't need to install any software; just open a browser to write and run code."
"The online environment also supports code highlighting, auto-completion, and error prompts.",
source="Online Programming Environment Description"
)
# You can also load local files; PDF and Markdown/TXT are automatically recognized:
# kb.add_file("./docs/产品手册.pdf")
# kb.add_file("./docs/常见问题.md")
# ========== Create a RAG Agent ==========
@tool
def search_knowledge(query: str) -> str:
"""Search the personal knowledge base for relevant information. Use the complete question or key phrases when searching.
Args:
query: the search question or key phrase
"""
docs = kb.search(query, k=3)
if not docs:
return "No relevant information found in the knowledge base."
results = []
for i, doc in enumerate(docs, 1):
source = doc.metadata.get("source", "Unknown source")
page = doc.metadata.get("page")
location = f"{source}" + (f" Page {page + 1}" if page is not None else "")
content = doc.page_content[:200]
results.append(f"[{i}] Source: {location}\n{content}")
return "\n\n---\n\n".join(results)
model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
agent = create_agent(
model=model,
tools=[search_knowledge],
system_prompt="""You are a personal knowledge base assistant.
## Rules
1. All questions must first use the search_knowledge tool to search the knowledge base
2. When answering, indicate the source of the information (document name, and page number if it is a PDF)
3. If there is no relevant content in the knowledge base, tell the user truthfully
4. Structure your answers using numbered lists or paragraphs""",
)
# ========== Test: non-streaming, and view the retrieved content at the same time ==========
def ask(question: str):
"""Ask a question and print both the retrieved raw snippets and the final answer for easy debugging"""
print(f"\n{'='*60}")
print(f"Q: {question}")
print(f"{'='*60}")
result = agent.invoke({
"messages": [HumanMessage(content=question)]
})
# Display the retrieved content
for msg in result["messages"]:
if msg.type == "tool":
print(f"\n[Retrieved content]")
print(msg.content[:300])
print(f"\n[Answer]")
print(result["messages"][-1].content)
# ========== Test: streaming, corresponding to "streaming output" in the system design ==========
def ask_stream(question: str):
"""Ask a question and print the answer token by token in a streaming manner, achieving a typewriter effect"""
print(f"\n{'='*60}")
print(f"Q: {question}")
print(f"{'='*60}")
print("\n[Answer] ", end="", flush=True)
for chunk, metadata in agent.stream(
{"messages": [HumanMessage(content=question)]},
stream_mode="messages",
):
# metadata["langgraph_node"] == "model" indicates that this chunk comes from the model-generated
# final answer node, filtering out other types of chunks such as tool calls
if metadata.get("langgraph_node") == "model" and chunk.content:
print(chunk.content, end="", flush=True)
print()
ask("What chapters does the Python3 basic tutorial contain?")
ask("How should I plan my Python learning path?")
ask_stream("What features does EXAMPLE's online programming environment support?")
Run result:
============================================================ Q: Python3 基础教程包含哪些章节? ============================================================ [检索到的内容] [1] 来源:Python3 教程大纲 Example 的 Python3 基础教程包含以下章节:... [回答] Python3 基础教程包含以下章节(来源:Python3 教程大纲): 1. Python 简介与环境搭建 2. 基本数据类型 3. 运算符与表达式 ... ============================================================ Q: 如何规划 Python 学习路线? ============================================================ [回答] 根据知识库中的 Python 学习路线建议(来源:Python 学习路线): 第一步:掌握基础语法(1-2 周) 第二步:学习数据结构和算法(2-3 周) 第三步:选择方向深入学习(Web/数据分析/AI) 第四步:做 2-3 个实战项目巩固 ============================================================ Q: Example的在线编程环境支持哪些功能? ============================================================ [回答] 根据知识库中的说明(来源:在线编程环境说明),Example的在线编程环境: 1. 支持 Python、JavaScript、Java、C++ 等多种语言 2. 无需安装任何软件,打开浏览器即可编写和运行代码 3. 支持代码高亮、自动补全和错误提示功能
Other ExtensionsThe last question uses
ask_stream()streaming output, so at runtime, the text after "[Answer]" is printed out one character (or a few characters per batch) at a time, just like a typewriter effect. Above, for convenient display, the final complete text after printing is pasted directly.