LangChain Personal Knowledge Base Q&A System

This section builds a personal knowledge base system that can load Markdown files and PDF documents and answer questions based on their content.


System Design

  • Document Loading: Supports multiple formats including Markdown, TXT, and PDF
  • Vector Retrieval: Chroma persistent storage, supports incremental updates
  • Source Citations: Answers include the source document and snippet location
  • Streaming Output: Displays the answer token by token

Complete Code

Before running, you need to configure the following.envin the file:DEEPSEEK_API_KEY(Chat model) andDASHSCOPE_API_KEY(Alibaba Cloud Bailian, used for knowledge base text vectorization). For specific application procedures and common troubleshooting, seethe "LangChain Intelligent Customer Service Robot"section.

Example

# File path: knowledge_qa.py
# pip install langchain langchain-deepseek langchain-openai langchain-community langchain-chroma chromadb pypdf
from dotenv import load_dotenv
load_dotenv()

import os
from pathlib import Path
from langchain.tools import tool
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage
from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import TextLoader, PyPDFLoader


class KnowledgeBase:
    """Personal knowledge base manager"""

    def __init__(self, persist_dir: str = "./my_knowledge_db"):
        self.persist_dir = persist_dir

        # Alibaba Cloud Bailian's Embedding interface is compatible with the OpenAI spec; just call it with langchain-openai
        # 's OpenAIEmbeddings — no need to install the DashScopeEmbeddings in langchain-community,
        # which has stopped being maintained.
        # check_embedding_ctx_length=False: turn off tiktoken pre-tokenization and send the original text directly
        # (Bailian's interface does not accept token id arrays).
        # chunk_size=10: Bailian's Embedding interface accepts at most 10 texts per request.
        self.embeddings = OpenAIEmbeddings(
            model="text-embedding-v4",
            api_key=os.getenv("DASHSCOPE_API_KEY"),
            base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
            check_embedding_ctx_length=False,
            chunk_size=10,
        )
        self.text_splitter = RecursiveCharacterTextSplitter(
            chunk_size=500, chunk_overlap=50,
            separators=["\n\n", "\n", "。", "!", "?", ". ", "! ", "? ", " "],
        )
        self.vector_store = None
        self._load_or_create()

    def _load_or_create(self):
        """Load an existing vector store or create a new one (Chroma works with the same set of parameters;
whether the directory already exists only affects the prompt message, not the creation logic)"""

        is_existing = os.path.exists(self.persist_dir) and os.listdir(self.persist_dir)

        self.vector_store = Chroma(
            embedding_function=self.embeddings,
            persist_directory=self.persist_dir,
        )

        if is_existing:
            print(f"Loaded vector store: {self.vector_store._collection.count()} document chunks")
        else:
            print("Created a new vector store")

    def add_file(self, file_path: str) -> int:
        """Add a file to the knowledge base and return the number of document chunks added.

Automatically choose the appropriate Loader based on the file extension:
- For .pdf use PyPDFLoader (depends on the pypdf library to parse PDFs, returns multiple Documents by page)
- For the rest (.md, .txt and other plain text files) use TextLoader to read the raw text
        """

        suffix = Path(file_path).suffix.lower()
        if suffix == ".pdf":
            loader = PyPDFLoader(file_path)
        else:
            loader = TextLoader(file_path, encoding="utf-8")

        docs = loader.load()

        # Uniformly add the file name as the source; each Document loaded by PyPDFLoader also comes with
        # a page field (page number), which can also be used to locate the snippet position during retrieval
        for doc in docs:
            doc.metadata["source"] = Path(file_path).name

        chunks = self.text_splitter.split_documents(docs)
        self.vector_store.add_documents(chunks)
        print(f"Added {Path(file_path).name}: {len(chunks)} document chunks")
        return len(chunks)

    def add_text(self, text: str, source: str = "Manually add") -> int:
        """Directly add text to the knowledge base"""
        chunks = self.text_splitter.create_documents(
            [text], metadatas=[{"source": source}]
        )
        self.vector_store.add_documents(chunks)
        return len(chunks)

    def search(self, query: str, k: int = 3) -> list:
        """Search the knowledge base"""
        return self.vector_store.similarity_search(query, k=k)

    def get_retriever(self):
        """Get the retriever"""
        return self.vector_store.as_retriever(search_kwargs={"k": 3})

Currently,TextLoader、PyPDFLoaderthis type of local file loader doesn't yet have a standalone maintained package; it can still only belangchain_community.document_loadersimported from langchain_community. At runtime, you may see a DeprecationWarning for the whole langchain-community package — this warning can be ignored for now; just migrate after the official standalone document loader package is released. This is different from the embedding case above, where a ready-made replacement already existed but wasn't adopted.

Continuing with the Agent section:

Example

# ========== Create the knowledge base and add sample data ==========

kb = KnowledgeBase("./my_knowledge_db")

# Add some sample knowledge
kb.add_text(
    "EXAMPLE's Python3 basic tutorial contains the following chapters:"
    "1. Python Introduction and Environment Setup 2. Basic Data Types 3. Operators and Expressions "
    "4. Conditional Statements if-else 5. Loops for/while 6. Function Definition and Calling "
    "7. Modules and Packages 8. File Operations 9. Exception Handling 10. Object-Oriented Programming",
    source="Python3 Tutorial Outline"
)

kb.add_text(
    "To become a good Python developer, it is recommended to study along the following route:"
    "First, master Python basic syntax (1-2 weeks);"
    "Second, learn the basics of data structures and algorithms (2-3 weeks);"
    "Third, choose a direction for in-depth study (Web Development/Data Analysis/AI);"
    "Fourth, work on 2-3 hands-on projects to consolidate your knowledge.",
    source="Python Learning Path"
)

kb.add_text(
    "EXAMPLE's online programming environment supports multiple languages including Python, JavaScript, Java, and C++."
    "Users don't need to install any software; just open a browser to write and run code."
    "The online environment also supports code highlighting, auto-completion, and error prompts.",
    source="Online Programming Environment Description"
)

# You can also load local files; PDF and Markdown/TXT are automatically recognized:
# kb.add_file("./docs/产品手册.pdf")
# kb.add_file("./docs/常见问题.md")


# ========== Create a RAG Agent ==========

@tool
def search_knowledge(query: str) -> str:
    """Search the personal knowledge base for relevant information. Use the complete question or key phrases when searching.

    Args:
query: the search question or key phrase
    """

    docs = kb.search(query, k=3)
    if not docs:
        return "No relevant information found in the knowledge base."

    results = []
    for i, doc in enumerate(docs, 1):
        source = doc.metadata.get("source", "Unknown source")
        page = doc.metadata.get("page")
        location = f"{source}" + (f" Page {page + 1}" if page is not None else "")
        content = doc.page_content[:200]
        results.append(f"[{i}] Source: {location}\n{content}")

    return "\n\n---\n\n".join(results)


model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
agent = create_agent(
    model=model,
    tools=[search_knowledge],
    system_prompt="""You are a personal knowledge base assistant.

## Rules
1. All questions must first use the search_knowledge tool to search the knowledge base
2. When answering, indicate the source of the information (document name, and page number if it is a PDF)
3. If there is no relevant content in the knowledge base, tell the user truthfully
4. Structure your answers using numbered lists or paragraphs"""
,
)


# ========== Test: non-streaming, and view the retrieved content at the same time ==========

def ask(question: str):
    """Ask a question and print both the retrieved raw snippets and the final answer for easy debugging"""
    print(f"\n{'='*60}")
    print(f"Q: {question}")
    print(f"{'='*60}")

    result = agent.invoke({
        "messages": [HumanMessage(content=question)]
    })

    # Display the retrieved content
    for msg in result["messages"]:
        if msg.type == "tool":
            print(f"\n[Retrieved content]")
            print(msg.content[:300])

    print(f"\n[Answer]")
    print(result["messages"][-1].content)


# ========== Test: streaming, corresponding to "streaming output" in the system design ==========

def ask_stream(question: str):
    """Ask a question and print the answer token by token in a streaming manner, achieving a typewriter effect"""
    print(f"\n{'='*60}")
    print(f"Q: {question}")
    print(f"{'='*60}")
    print("\n[Answer] ", end="", flush=True)

    for chunk, metadata in agent.stream(
        {"messages": [HumanMessage(content=question)]},
        stream_mode="messages",
    ):
        # metadata["langgraph_node"] == "model" indicates that this chunk comes from the model-generated
        # final answer node, filtering out other types of chunks such as tool calls
        if metadata.get("langgraph_node") == "model" and chunk.content:
            print(chunk.content, end="", flush=True)
    print()


ask("What chapters does the Python3 basic tutorial contain?")
ask("How should I plan my Python learning path?")
ask_stream("What features does EXAMPLE's online programming environment support?")

Run result:

============================================================
Q: Python3 基础教程包含哪些章节?
============================================================

[检索到的内容]
[1] 来源:Python3 教程大纲
Example 的 Python3 基础教程包含以下章节:...

[回答]
Python3 基础教程包含以下章节(来源:Python3 教程大纲):
1. Python 简介与环境搭建
2. 基本数据类型
3. 运算符与表达式
...

============================================================
Q: 如何规划 Python 学习路线?
============================================================

[回答]
根据知识库中的 Python 学习路线建议(来源:Python 学习路线):
第一步:掌握基础语法(1-2 周)
第二步:学习数据结构和算法(2-3 周)
第三步:选择方向深入学习(Web/数据分析/AI)
第四步:做 2-3 个实战项目巩固

============================================================
Q: Example的在线编程环境支持哪些功能?
============================================================

[回答] 根据知识库中的说明(来源:在线编程环境说明),Example的在线编程环境:
1. 支持 Python、JavaScript、Java、C++ 等多种语言
2. 无需安装任何软件,打开浏览器即可编写和运行代码
3. 支持代码高亮、自动补全和错误提示功能

The last question usesask_stream()streaming output, so at runtime, the text after "[Answer]" is printed out one character (or a few characters per batch) at a time, just like a typewriter effect. Above, for convenient display, the final complete text after printing is pasted directly.

Other Extensions