☰
第三单元——第七课:把向量库保存到磁盘
2026/10/5 4:54:31 网站建设 项目流程

之前每次运行都会重新计算bookstore.txt的文档向量。这次用 Chroma 保存索引:首次运行建立,之后运行直接读取。LangChain 的 Chroma 文档

先在虚拟环境中安装:

python -m pip install -U langchain-chroma

保存为unit3_lesson7_persistent_index.py:

import time from pathlib import Path import chromadb from langchain_chroma import Chroma from langchain_core.documents import Document from langchain_huggingface import HuggingFaceEmbeddings from langchain_text_splitters import RecursiveCharacterTextSplitter start = time.perf_counter() folder = Path(__file__).parent database_path = folder / "bookstore_chroma" collection_name = "bookstore_lessons" # 本地模型仍需加载,用来把新问题转换成向量 embeddings = HuggingFaceEmbeddings( model_name="BAAI/bge-small-zh-v1.5", model_kwargs={"device": "cpu", "local_files_only": True}, encode_kwargs={"normalize_embeddings": True}, ) client = chromadb.PersistentClient(path=str(database_path)) collection = client.get_or_create_collection(collection_name) vector_store = Chroma( client=client, collection_name=collection_name, embedding_function=embeddings, ) # 只有空索引才读取文件、切分并计算文档向量 if collection.count() == 0: file_path = folder / "bookstore.txt" document = Document( page_content=file_path.read_text(encoding="utf-8"), metadata={"source": file_path.name}, ) splitter = RecursiveCharacterTextSplitter( chunk_size=120, chunk_overlap=20, add_start_index=True, separators=["\n\n", "\n", "。", ",", ""], ) chunks = splitter.split_documents([document]) vector_store.add_documents(chunks) print(f"首次建立索引:{len(chunks)} 个片段") else: print(f"复用已有索引:{collection.count()} 个片段") question = "会员买书有什么优惠?" results = vector_store.similarity_search(question, k=1) print("检索结果:") for doc in results: print(doc.page_content) print("来源:", doc.metadata["source"]) print("起始位置:", doc.metadata["start_index"]) print(f"总耗时:{time.perf_counter() - start:.2f} 秒")

连续运行两次:

python unit3_lesson7_persistent_index.py python unit3_lesson7_persistent_index.py

第一次应显示“首次建立索引”,第二次显示“复用已有索引”。磁盘上会出现bookstore_chroma文件夹。

这里节省的是重复计算文档向量的时间。本地 Embedding 模型每次启动仍需加载,而且问题本身仍需转换成向量;由于示例文档很短,两次总耗时可能差别不大。

本课代码只在索引为空时写入。如果之后修改了bookstore.txt,旧索引不会自动更新;这正是下一课要解决的问题。

课后问题:bookstore_chroma 是通过执行哪行代码后创建的?

回答:主要是这一行创建bookstore_chroma文件夹:

client = chromadb.PersistentClient(path=str(database_path))

前一行database_path = folder / "bookstore_chroma"只是指定路径;PersistentClient初始化时会在该位置建立持久化数据库。首次运行中的vector_store.add_documents(chunks)则把文档片段和向量写入数据库。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询