- AI Agent
- AI 应用
- 后端
- 前端
- 大模型
- RAG
【免费下载链接】nexent
Nexent is a zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles — unified tools, skills, memory, and orchestration with built-in constraints, feedback loops, and control planes.
本文是一份围绕 Nexent 开源项目(零代码 AI Agent 自动生成平台)中 SDK 向量数据库模块的实战技术指南。它完整讲解nexent.vector_database提供的 Elasticsearch 向量检索与文档管理服务:如何通过 Docker 部署 Elasticsearch 8.17.4 与 Kibana、如何通过EmbeddingAdapter网关适配器(Jina、OpenAI 兼容、DashScope、SiliconFlow 等)为文档生成 embedding 向量,以及如何使用ElasticSearchCore完成索引管理、批量向量化入库、精确检索、语义检索与混合检索。读完本文,你将掌握在本地或远程环境完整落地一套基于 Elasticsearch 的知识库检索后端,并理解其统一向量存储抽象与 REST API 边界。
一、模块定位:统一的向量存储抽象
Nexent SDK 的向量数据库能力集中在sdk/nexent/vector_database/目录,核心设计是"一套统一接口、多种后端实现":
- sdk/nexent/vector_database/base.py:定义抽象基类
VectorDatabaseCore(继承ABC),统一声明索引管理、文档操作、检索操作、统计监控四大类抽象方法。服务层只依赖该抽象接口,因此后续可以平滑扩展到 Milvus 等其他向量数据库后端,而业务代码无需改动。 - sdk/nexent/vector_database/elasticsearch_core.py:
ElasticSearchCore,本指南的主体,封装了所有 Elasticsearch 操作。 - sdk/nexent/vector_database/datamate_core.py:
DataMateCore,基于 DataMate REST API 的向量存储实现。注意nexent.vector_database包的默认导出(__all__ = ["DataMateCore"])是它,而本文讲解的ElasticSearchCore需要从nexent.vector_database.elasticsearch_core显式导入。 - sdk/nexent/vector_database/utils.py:数据格式化与查询 DSL 构建工具,如
format_size、build_weighted_query。
embedding 向量统一由nexent.core.gateway.modality下的EmbeddingAdapter适配器生成,包括JinaEmbeddingAdapter、OpenAICompatibleEmbeddingAdapter、DashScopeEmbeddingAdapter、SiliconflowEmbeddingAdapter等。从源码看,前三者继承自_MultimodalEmbeddingAdapter(支持文本与图像多模态输入),OpenAICompatibleEmbeddingAdapter则直接继承EmbeddingAdapter,通过 OpenAI 兼容的/embeddings接口生成文本向量,实现见 sdk/nexent/core/gateway/modality/embedding/openai.py。
二、环境准备与基础初始化
1. 安装依赖
向量数据库核心依赖官方 Elasticsearch Python 客户端,安装命令如下:
pip install elasticsearch2. 连接凭据通过构造参数传入
需要特别强调的是:SDK 本身不读取任何环境变量,连接凭据全部通过构造函数参数显式传入;环境变量由服务层读取后再传递给 SDK。
vdb_core = ElasticSearchCore( host="https://localhost:9200", api_key="your_api_key", )embedding 模型以EmbeddingAdapter适配器形式传入,例如JinaEmbeddingAdapter、OpenAICompatibleEmbeddingAdapter等。构造函数内部(见 elasticsearch_core.py)会使用这两个参数创建Elasticsearch客户端,并预设request_timeout=20、max_retries=3、retry_on_timeout=True、retry_on_status=[502, 503, 504],同时在客户端内部定义 embedding API 的批处理上限:单批最多 2048 条文本、单条文本最多 8192 token、总 token 上限 100000。
三、Docker 部署指南
1. 前置条件
- 安装 Docker(访问 Docker 官网下载安装)。
- 若使用 Docker Desktop,请为其分配至少 4GB 内存,可在Settings > Resources中调整。
- 创建 Docker 网络:
docker network create elastic2. Elasticsearch 部署
- 拉取镜像:
docker pull docker.elastic.co/elasticsearch/elasticsearch:8.17.4- 启动容器(后台模式,首次启动需等待 3-5 分钟):
docker run -d --name es01 --net elastic -p 9200:9200 -m 6GB -e "xpack.ml.use_auto_machine_memory_percent=true" docker.elastic.co/elasticsearch/elasticsearch:8.17.4- 查看启动日志:
docker logs -f es01- 重置
elastic用户密码(确认时输入 Yes):
docker exec -it es01 /usr/share/elasticsearch/bin/elasticsearch-reset-password -u elastic- 保存重要信息:容器启动时会显示
elastic用户密码与 Kibana enrollment token,建议将密码保存为环境变量:
export ELASTIC_PASSWORD="your_password"- 拷贝 SSL 证书(后续 curl 校验与 SDK 连接都需要):
docker cp es01:/usr/share/elasticsearch/config/certs/http_ca.crt .- 验证部署:
curl --cacert http_ca.crt -u elastic:$ELASTIC_PASSWORD https://localhost:9200 -k- 获取 API key(SDK 连接推荐使用 API key 而非账号密码):
curl --cacert http_ca.crt \ -u elastic:$ELASTIC_PASSWORD \ --request POST \ --url https://localhost:9200/_security/api_key \ --header 'Content-Type: application/json' \ --data '{ "name": "pick-a-name" }'- 验证 API key 可用:
curl --request GET \ --url https://XXX.XX.XXX.XX:9200/_cluster/health \ --header 'Authorization: ApiKey API-KEY'3. Kibana 部署(可选)
- 拉取 Kibana 镜像并启动(版本需与 Elasticsearch 一致,均为 8.17.4):
docker pull docker.elastic.co/kibana/kibana:8.17.4 docker run -d --name kib01 --net elastic -p 5601:5601 docker.elastic.co/kibana/kibana:8.17.4- 查看 Kibana 日志:
docker logs -f kib01- 配置 Kibana:在 es01 容器内生成 enrollment token,然后在浏览器访问 http://localhost:5601 并输入该 token:
docker exec -it es01 /usr/share/elasticsearch/bin/elasticsearch-create-enrollment-token -s kibana提示:页面可能需要输入验证码,可通过
docker logs -f kib01查看。
- 使用
elastic用户和前面生成的密码登录 Kibana。
4. 常见管理命令
# 停止容器 docker stop es01 docker stop kib01 # 删除容器 docker rm es01 docker rm kib01 # 删除网络 docker network rm elastic5. 生产环境注意事项
- 数据持久化:必须将数据卷挂载到
/usr/share/elasticsearch/data,否则容器删除后数据全部丢失。示例启动命令:
docker run -d --name es01 --net elastic -p 9200:9200 -m 6GB -v es_data:/usr/share/elasticsearch/data docker.elastic.co/elasticsearch/elasticsearch:8.17.4内存配置:按实际需要调整容器内存上限,建议至少 6GB。
常见故障排查:
- 内存不足:检查 Docker Desktop 的内存设置;
- 端口冲突:确认 9200 端口未被占用;
- 证书问题:确认 SSL 证书拷贝正确;
- 昇腾/物理机部署时的
vm.max_map_count问题:ES 启动会因内核参数过小而失败,报错max virtual memory areas vm.max_map_count [65530] is too low, increase to at least [262144]。需要在宿主机上临时或持久化调整:
# 临时生效(在宿主机执行) sudo sysctl -w vm.max_map_count=262144 # 持久化:编辑 /etc/sysctl.conf 添加 vm.max_map_count=262144 # 然后执行 sudo sysctl -p6. 远程部署排查指南
当 Elasticsearch 部署在远程服务器时,常遇到网络访问问题,常见问题与解决方案如下:
- 远程访问被拒绝:curl 返回 "Connection reset by peer"。可通过 SSH 隧道做端口转发,再通过本地端口访问:
# 使用 SSH 隧道转发端口 ssh -L 9200:localhost:9200 user@remote_server # 在新终端通过本地端口访问 curl -H "Authorization: ApiKey your_api_key" https://localhost:9200/_cluster/health?pretty -k- 网络配置检查清单:
- 确保远程服务器防火墙放行 9200 端口(以 iptables 为例):
sudo iptables -A INPUT -p tcp --dport 9200 -j ACCEPT sudo service iptables save- 检查 Elasticsearch 网络配置(elasticsearch.yml 示例):
network.host: 0.0.0.0 http.cors.enabled: true http.cors.allow-origin: "*"安全配置建议(生产环境):
- 将 CORS
allow-origin限制为具体域名; - 使用反向代理(如 Nginx)管理 SSL 终止;
- 配置合适的网络安全组规则;
- 使用受信任的 SSL 证书而非自签名证书。
- 将 CORS
使用环境变量管理远程连接:在
.env文件中配置:
ELASTICSEARCH_HOST=https://remote_server:9200 ELASTICSEARCH_API_KEY=your_api_key如果使用 SSH 隧道,可继续使用 localhost:
ELASTICSEARCH_HOST=https://localhost:9200- 故障排查命令:
# 检查端口监听状态 netstat -tulpn | grep 9200 # 查看 ES 日志 docker logs es01 # 测试 SSL 连接 openssl s_client -connect remote_server:9200四、核心组件解析
| 文件 | 核心类 | 职责 |
|---|---|---|
elasticsearch_core.py | ElasticSearchCore | 全部 Elasticsearch 操作:索引管理、向量化入库、三类检索、统计监控 |
base.py | VectorDatabaseCore | 抽象基类,定义统一的向量存储接口,便于扩展其他后端 |
datamate_core.py | DataMateCore | DataMate 向量存储实现(nexent.vector_database包的默认导出) |
utils.py | format_size/build_weighted_query | 数据格式化与查询 DSL 构建 |
从源码看,VectorDatabaseCore抽象接口按功能划分为四组:索引管理(create_index/delete_index/get_user_indices/check_index_exists)、文档操作(vectorize_documents/delete_documents/get_index_chunks/create_chunk/update_chunk/delete_chunk/count_documents)、检索操作(accurate_search/semantic_search/hybrid_search)、统计监控(get_documents_detail/get_indices_detail)。DataMateCore因受 DataMate HTTP API 能力限制,对部分操作显式抛出NotImplementedError以明确边界(详见 datamate_core.py)。
embedding 向量由nexent.core.gateway.modality下的EmbeddingAdapter适配器生成(例如JinaEmbeddingAdapter、OpenAICompatibleEmbeddingAdapter、DashScopeEmbeddingAdapter、SiliconflowEmbeddingAdapter)。以OpenAICompatibleEmbeddingAdapter为例,其get_embeddings将输入归一化为列表后调用 OpenAI 兼容的/embeddings接口,并对超时进行指数退避重试(retries=3、每次递增retry_timeout_step=5.0秒),实现见 openai.py。
五、使用示例
1. 基础初始化
from nexent.vector_database.elasticsearch_core import ElasticSearchCore # 直接指定凭据(host 与 api_key 为必填参数) vdb_core = ElasticSearchCore( host="https://localhost:9200", api_key="your_api_key", verify_certs=False, ssl_show_warn=False, )verify_certs=False适用于自签名证书的本地开发环境;生产环境应替换为受信任证书并开启校验。SDK 客户端还内置request_timeout=20、max_retries=3等稳健性参数,无需外部配置。
2. 索引管理
# 创建新的向量索引(embedding_dim 可选;不指定时使用 embedding 模型的维度) vdb_core.create_index("my_documents") # 列出所有用户索引 indices = vdb_core.get_user_indices() print(indices) # 检查索引是否存在 exists = vdb_core.check_index_exists("my_documents") print(exists) # 删除索引 vdb_core.delete_index("my_documents")源码细节(elasticsearch_core.py):create_index默认按embedding_dim or 1024创建dense_vector字段(index=true、similarity=cosine),索引设置采用固定均衡参数(1 分片、0 副本、refresh_interval=5s、max_result_window=50000、异步 translog 等),并自动跳过已存在的索引、强制 refresh 并等待集群状态变黄。
3. 文档操作
from nexent.core.gateway.model_context import EmbeddingContext from nexent.core.gateway.modality import OpenAICompatibleEmbeddingAdapter # 构建 embedding 模型适配器 embedding_model = OpenAICompatibleEmbeddingAdapter(EmbeddingContext( model_name="your-embedding-model", base_url="https://your-embedding-api/v1/embeddings", api_key="your_api_key", modality="embedding", factory="openai", embedding_dim=1024, ))EmbeddingContext是ModelContext的 dataclass 子类(见 model_context.py),额外提供embedding_dim(向量维度)与model_type("embedding"或"multi_embedding",对应纯文本与多模态)两个字段,用于描述 embedding 模型的上下文。
# 批量索引文档(embedding 向量自动生成;batch_size 默认 64) documents = [ { "id": "doc1", "title": "Document 1", "file": "file1.txt", "path_or_url": "https://example.com/doc1", "content": "This is the content of document 1", "process_source": "Web", "embedding_model_name": "your-embedding-model", # 指定 embedding 模型 "file_size": 1024, # 文件大小(字节) "create_time": "2023-06-01T10:30:00" # 文件创建时间 }, { "id": "doc2", "title": "Document 2", "file": "file2.txt", "path_or_url": "https://example.com/doc2", "content": "This is the content of document 2", "process_source": "Web" # 未提供的字段使用默认值 } ] total_indexed = vdb_core.vectorize_documents( "my_documents", embedding_model, documents, batch_size=64 ) print(f"Successfully indexed {total_indexed} documents") # 按 URL 或路径删除文档 deleted_count = vdb_core.delete_documents("my_documents", "https://example.com/doc1") print(f"Deleted {deleted_count} documents")向量化入库的智能策略(源码级原理,elasticsearch_core.py):vectorize_documents会根据数据量自动选择策略:
- 大批量路径(
total_docs >= 64或large_mode=True):进入bulk_operation_context上下文管理器,先把索引的refresh_interval调整为 30s、translog 改为异步(_apply_bulk_settings),批量写入完成后恢复refresh_interval=5s、translog.durability=request(_restore_normal_settings),显著提升写入吞吐。 - 小批量路径:直接插入,并以
refresh="wait_for"保证实时可见。
两条路径都会把文档按embedding_batch_size(默认 10)拆分为更小的子批调用 embedding API,并带最多 3 次指数退避重试,避免单个 provider 瞬时故障拖垮整批写入。_preprocess_documents会自动补齐缺失字段:create_time默认当前时间、date默认当前日期、file_size默认 0、process_source默认"Unstructured"、id默认按时间戳与内容 hash 生成。此外,若 embedding 模型不是多模态类型,process_source == "UniversalImageExtractor"的图像块会被过滤;多模态模型则走get_multimodal_embeddings,图像块向量写入multi_embedding字段。delete_documents底层通过delete_by_query以term查询path_or_url精确删除。
4. 检索
# 精确文本检索(index_names 为索引名列表,支持多索引) results = vdb_core.accurate_search(["my_documents"], "sample query", top_k=5) for result in results: print(f"Score: {result['score']}, Document: {result['document']['title']}") # 语义向量检索(必须传入 embedding 模型) results = vdb_core.semantic_search(["my_documents"], "sample query", embedding_model, top_k=5) for result in results: print(f"Score: {result['score']}, Document: {result['document']['title']}") # 混合检索(weight_accurate 可选;默认自动推断:含数字的查询偏向精确检索 0.7,否则 0.3) results = vdb_core.hybrid_search( ["my_documents"], "sample query", embedding_model, top_k=5, weight_accurate=0.3 # 精确检索权重 0.3,向量检索权重 0.7 ) for result in results: print(f"Score: {result['score']}, Document: {result['document']['title']}")三种检索的源码实现(elasticsearch_core.py):
accurate_search:调用build_weighted_query(见 utils.py)构造加权查询 DSL——对每个词项按term_weights * field_weights * boost_factor(boost_factor 默认 2.0)生成function_score权重函数,同时用match_phrase(slop=3)与match(minimum_should_match=50%、fuzziness=AUTO)在title、content字段上做模糊匹配,默认field_weights={"title": 1, "content": 1}。检索结果会携带高亮片段,并通过内部标记__nexent_hit_start__/__nexent_hit_end__提取命中词供结果卡片解释使用。它还支持可选的 ESfilter子句(如{"bool": {"must": [...]}}),用于内存索引隔离等场景,作为bool.filter与打分查询取交集。semantic_search:先用 embedding 模型对查询文本向量化,再对embedding字段执行 kNN 检索(k=top_k、num_candidates=top_k*2,可选 filter 作用于候选选择阶段);多模态模型还会同时对multi_embedding字段做一次 kNN,并合并两次结果。hybrid_search:融合精确与语义两路结果。先按文档id归并两侧分数,再分别除以各自批次最大值做归一化,最后按combined_score = weight_accurate * normalized_accurate + (1 - weight_accurate) * normalized_semantic加权求和并降序返回。当调用方未指定weight_accurate时,含数字的查询(告警编号、IP 等)自动采用 0.7,其余采用 0.3。
5. 统计与监控
# 获取索引统计信息 stats = vdb_core.get_indices_detail(["my_documents"]) print(stats) # 获取文件列表详情 file_details = vdb_core.get_documents_detail("my_documents") print(file_details) # 统计索引内文档数量 doc_count = vdb_core.count_documents("my_documents") print(doc_count) # 分页拉取索引文本块 chunks = vdb_core.get_index_chunks("my_documents", page=1, page_size=10) print(chunks)六、ElasticSearchCore 主要功能与高级特性
ElasticSearchCore提供以下核心能力:
- 索引管理:创建/删除索引、检查索引是否存在、列出用户索引;
- 文档操作:带 embedding 向量的批量文档索引、按
path_or_url删除指定文档、单块级 CRUD(create_chunk/update_chunk/delete_chunk); - 检索操作:精确文本检索、语义向量检索、混合检索(均支持可选 ES filter);
- 统计与监控:索引统计(
get_indices_detail)、文件列表详情(get_documents_detail)、文档计数(count_documents)。
高级特性示例
# 获取文件列表详情(返回字段:path_or_url, filename, file_size, create_time) files = vdb_core.get_documents_detail("my_documents") for file in files: print(f"File path: {file['path_or_url']}") print(f"Filename: {file['filename']}") print(f"File size: {file['file_size']} bytes") print(f"Created at: {file['create_time']}") print("---") # 获取全部索引的聚合统计 all_stats = vdb_core.get_indices_detail(["my_documents", "other_index"]) for index_name, stats in all_stats.items(): print(f"Index: {index_name}") print(f"Document count: {stats['base_info']['doc_count']}") print(f"Embedding model: {stats['base_info'].get('embedding_model')}") print("---")补充源码细节:get_documents_detail通过terms聚合按path_or_url去重(最多 1000 个文件),并用top_hits取每个文件的样本字段(path_or_url、file_size、create_time、filename),同时返回chunk_count(该文件的分块数);get_indices_detail则组合索引 stats、settings 与多重聚合(path_or_url基数、process_source分布、embedding_model_name分布),返回base_info(doc_count、chunk_count、store_size、process_source、embedding_model、embedding_dim、creation_date、update_date)与search_performance(total_search_count、hit_count)两组指标。
get_index_chunks支持两种模式:传入page+page_size时走from/size分页;不传时使用 ES scroll(TTL 2 分钟、每批 1000 条)全量拉取,并在finally中清理 scroll 上下文。单块 CRUD 中,update_chunk/delete_chunk会先尝试把传入的chunk_id当作 ES_id,失败后回退到按id字段的term查询定位(_resolve_chunk_document_id),从而兼容两种 id 约定。
七、REST API 边界说明
SDK 本身不附带 REST 服务。当前仓库中,知识库相关的 REST API 由后端服务提供,路由定义在 backend/apps/northbound_knowledge_app.py,主要包括:
- POST
/indices/{index_name}:创建索引; - DELETE
/indices/{index_name}:删除索引; - POST
/indices/search/hybrid:混合检索(内部把知识库名称解析为真实索引名,再调用ElasticSearchService.search_hybrid,见 northbound_knowledge_app.py); - DELETE
/indices/{index_name}/documents?path_or_url=...&scope=...:删除文档scope=source_only:仅删除 MinIO 源文件,保留 ES 中的分块与向量(检索仍可用,预览不可用);scope=full:删除 ES 文档与 MinIO 源文件,并清理相关 Redis 任务记录(删除接口在scope=full时会联动redis_service.delete_document_records清理 Celery 任务与缓存键,见 northbound_knowledge_app.py);
- GET
/indices/{index_name}/files:获取索引文件列表。
请求/响应的精确格式以后端路由定义为准。后端对该抽象接口的消费与校验在 test/backend/services/test_vectordatabase_service.py 中有系统覆盖(例如test_accurate_search、test_semantic_search、test_search_hybrid_success、test_search_hybrid_weight_accurate_boundary_values等用例),可作为接入行为的参考验证。
八、License
本项目基于 MIT License 开源,详见仓库根目录 LICENSE。
- AI Agent
- AI 应用
- 后端
- 前端
- 大模型
- RAG
【免费下载链接】nexent
Nexent is a zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles — unified tools, skills, memory, and orchestration with built-in constraints, feedback loops, and control planes.
相关推荐
langchaingo 接入 Chroma 向量数据库实战指南:从 Docker 部署到相似度检索与元数据过滤
langchaingo 接入 Chroma 向量数据库实战指南:从 Docker 部署到相似度检索与元数据过滤 本篇技术指南围绕 langchaingo 内置的
人工智能大模型AI AgentRAG后端Elasticsearch Setup 实战指南:将 NLWeb 检索后端接入 Elasticsearch 向量数据库
Elasticsearch Setup 实战指南:将 NLWeb 检索后端接入 Elasticsearch 向量数据库 导读 本文面向 NLWeb(Python
AI 应用MCP 服务AI Agent后端前端Milvus 向量数据库实战:从 Docker 部署到 Visualized-BGE 图文多模态检索(all-in-rag 系列)
Milvus 向量数据库实战:从 Docker 部署到 Visualized BGE 图文多模态检索(all in rag 系列) 本文是 Datawhale
示例工程教程人工智能大模型RAG
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考