1. 为什么 Deepseek API 原生工具调用总报错
你手里有一套跑得挺顺的 Agent 框架,LangChain、Cline、或者自己写的 OpenAI 兼容客户端,本来接 GPT 系列一切正常。换成 Deepseek API 之后,只要一传tools参数,要么直接 400,要么模型把工具描述当普通文本念出来,finish_reason永远是stop,tool_calls字段压根不出现。这个场景我见得太多了,核心原因就一句话:Deepseek 官方 API 的 chat completions 接口在工具调用这块的兼容层并不完整,很多部署版本只认messages和temperature这类基础字段,tools、tool_choice、parallel_tool_calls传过去会被静默忽略或者直接拒绝。
更麻烦的是,Agent 框架内部对工具调用的期待是 OpenAI 那套标准协议:模型返回finish_reason: "tool_calls",message.tool_calls是一个数组,每个元素带id、type: "function"、function.name和function.arguments。框架拿到这个结构才会去执行本地函数,把结果以role: "tool"的消息塞回对话。Deepseek 原生不吐这个结构,整条链路就断了。
那能不能不改框架代码?能。思路是在中间加一层协议适配,把 OpenAI 标准的tools请求翻译成 Deepseek 能理解的 System Prompt,再把 Deepseek 返回的文本里解析出工具调用意图,重新包装成标准tool_calls响应。LiteLLM 的 custom callback 机制正好能干这件事,而 TaoToken 提供的统一 Key 和 API 通道,让你不用为每个模型单独配 base_url 和鉴权,一个 Key 打通所有后端。
这篇要解决的就是:Deepseek API 工具调用不兼容时,怎么用 TaoToken 统一 Key 接入 LiteLLM,让 Agent 框架零改造直接跑 function calling。适合正在用 LangChain、Cline、Continue 这类工具,又想把后端换成 Deepseek 的开发者。下面从拿 Key 开始,到 config.yaml 配置,再到实际发一个带 tools 的请求验证返回结构,一步步来。
2. TaoToken 统一 Key 与 LiteLLM 接入前置准备
先说清楚 TaoToken 在这套方案里的角色。它提供的是一个 OpenAI 兼容的统一 API 入口,你拿一个 Key,就能通过同一个 Base URL 访问包括 Deepseek 在内的多种模型。对 LiteLLM 来说,这意味着api_base只需要填 TaoToken 的地址,api_key也只需要填 TaoToken 的 Key,不用为每个模型维护不同的凭证。模型 ID 通过model字段区分,LiteLLM 会把它透传到 TaoToken 的/v1/chat/completions。
你需要准备的东西不多:一个 TaoToken 账号、一个 API Key、本地装好 Python 3.9+ 和 pip。LiteLLM 的安装直接用 pip 就行,它自带 proxy server 模式,我们后面要用litellm --config启动。
先拿 Key。打开 TaoToken 控制台,在 API Keys 页面创建一个新 Key,复制出来。这个 Key 的权限范围覆盖你账号下可用的模型,Deepseek 系列也在里面。拿到之后先别急着写代码,用 curl 测一下通道是否通:
curl https://taotoken.net/api/v1/chat/completions \ -H "Authorization: Bearer sk-你的TaoTokenKey" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek-chat", "messages": [{"role": "user", "content": "你好"}] }'如果返回正常的choices结构,说明 Key 和通道都没问题。注意这里的 Base URL 是https://taotoken.net/api,后面 LiteLLM 配置里的api_base要写成https://taotoken.net/api/v1,因为 LiteLLM 会在后面拼/chat/completions。这个细节踩过坑,写错了会 404。
接下来装 LiteLLM。建议用虚拟环境,避免和系统 Python 包冲突:
python -m venv litellm-env source litellm-env/bin/activate pip install "litellm[proxy]" jinja2jinja2是后面写工具描述模板要用的,LiteLLM 的 callback 里会渲染 System Prompt。装完之后litellm --version能输出版本号就 OK。
这里要强调一个前置认知:LiteLLM 的 custom callback 机制允许你在请求进入模型之前和响应返回之后插入处理逻辑。我们要做的工具调用适配,就是在async_pre_call_hook里把tools转成文本塞进 System Prompt,在async_post_call_success_hook里把模型返回的 JSON 文本解析回tool_calls结构。TaoToken 负责把请求稳定送到 Deepseek 后端,LiteLLM 负责协议转换,两边各司其职。
3. 可复制的 LiteLLM config.yaml 与工具适配配置
这一节是核心,直接给能跑的配置和代码。先建目录结构,建议这样组织:
deepseek-tools/ ├── config.yaml ├── tool_handler.py └── templates/ └── prompt_template.j2config.yaml里定义模型列表和 callback 挂载点。关键在model_info.use_proxy: true,这个标记决定哪些模型走工具适配逻辑,没标记的模型保持原生行为:
model_list: - model_name: deepseek-chat litellm_params: model: openai/deepseek-chat api_base: https://taotoken.net/api/v1 api_key: os.environ/TAOTOKEN_API_KEY model_info: use_proxy: true - model_name: deepseek-reasoner litellm_params: model: openai/deepseek-reasoner api_base: https://taotoken.net/api/v1 api_key: os.environ/TAOTOKEN_API_KEY model_info: use_proxy: true litellm_settings: callbacks: tool_handler.tool_call_handler_instance注意api_key用os.environ/TAOTOKEN_API_KEY从环境变量读,别硬编码在文件里。启动前export TAOTOKEN_API_KEY=sk-你的Key。model字段前缀openai/是告诉 LiteLLM 用 OpenAI 兼容协议发请求,TaoToken 的入口正好是这个协议。
然后是tool_handler.py,这是适配逻辑的主体。我把它精简到能直接跑的程度,保留了核心的请求拦截和响应解析:
import json import os from typing import Literal, Optional from jinja2 import Environment, FileSystemLoader, select_autoescape from litellm.integrations.custom_logger import CustomLogger from litellm.proxy.proxy_server import UserAPIKeyAuth, DualCache, proxy_config from litellm.types.utils import ModelResponse, ChatCompletionMessageToolCall from litellm.caching.caching import Cache from litellm._logging import verbose_proxy_logger class ToolHandler(CustomLogger): def __init__(self, message_logging=True): super().__init__(message_logging) template_path = os.path.join(os.path.dirname(__file__), "templates") self.env = Environment( loader=FileSystemLoader(template_path), autoescape=select_autoescape(), ) self.function_cache = Cache(type="local", default_in_memory_ttl=180.0) def support_proxy(self, data: dict) -> bool: model_list = proxy_config.config["model_list"] for model in model_list: if model["model_name"] == data["model"]: if "use_proxy" in model["model_info"]: return model["model_info"]["use_proxy"] return False async def async_pre_call_hook( self, user_api_key_dict: UserAPIKeyAuth, cache: DualCache, data: dict, call_type: Literal["completion", "text_completion", "embeddings", "image_generation", "moderation", "audio_transcription"], ): if not self.support_proxy(data) or call_type != "completion": return data if "tools" in data and isinstance(data["tools"], list) and len(data["tools"]) > 0: template = self.env.get_template("prompt_template.j2") rendered_prompt = template.render(tools=data["tools"]) if "messages" in data: for message in data["messages"]: if message.get("role") == "system": message["content"] = f"{message['content']}\n\n{rendered_prompt}" data["tools"] = None if "messages" in data: for message in data["messages"]: if "tool_calls" in message: del message["tool_calls"] if message.get("role") == "tool": tool_call_id = message.get("tool_call_id") if tool_call_id is not None: tool_call = await self.function_cache.async_get_cache(cache_key=tool_call_id) if tool_call is not None and isinstance(tool_call, dict): function_name = tool_call["function"]["name"] message["role"] = "user" content = message["content"] message["content"] = f"Function: {function_name}, Result: {content}" del message["tool_call_id"] return data async def async_post_call_success_hook( self, data: dict, user_api_key_dict: UserAPIKeyAuth, response: ModelResponse, ): if not self.support_proxy(data): return choice = response.choices[0] content = choice.message.content tool_calls = [] try: json_blocks = content.split("```json") for block in json_blocks[1:]: json_str = block.split("```")[0].strip() json_data = json.loads(json_str) if isinstance(json_data, list): for item in json_data: if "name" in item and "arguments" in item: tool_call = ChatCompletionMessageToolCall(function=item) tool_calls.append(tool_call) await self.function_cache.async_add_cache(tool_call, cache_key=tool_call.id) else: if "name" in json_data and "arguments" in json_data: tool_call = ChatCompletionMessageToolCall(function=json_data) tool_calls.append(tool_call) await self.function_cache.async_add_cache(tool_call, cache_key=tool_call.id) if len(tool_calls) > 0: choice.finish_reason = "tool_calls" choice.message.tool_calls = tool_calls response.choices = [choice] except Exception as e: verbose_proxy_logger.error(f"Error parsing JSON Markdown: {e}") tool_call_handler_instance = ToolHandler()templates/prompt_template.j2负责把工具定义渲染成模型能读懂的指令:
{% if tools|length > 0 %} 你可以用如下工具解决问题: {% for tool in tools %}{% if tool['type'] == 'function' %} ### {{ tool['function']['name'] }} ```json {{ tool['function'] | tojson(indent=4) }}{% endif %}{% endfor %}
注意:
- 在调用上述函数时,请使用 Json 格式表示调用的参数。
- 用尽量少的函数调用解决问题。
- 调用JSON格式参考如下:
{ "name": "send_email", "arguments": { "to_email": "Recipient's email address", "title": "The title of the email", "body": "The Body of the email" } }{% endif %}
这套配置的逻辑是:请求进来时,`tools` 被渲染成 System Prompt 的一部分,原始 `tools` 字段置空,Deepseek 看到的就是一段带工具说明的普通对话。模型按模板要求输出 JSON 代码块,响应回来时 callback 解析代码块,构造出标准 `tool_calls` 结构,同时把 `finish_reason` 改成 `tool_calls`。Agent 框架收到的东西和调 GPT 时一模一样。 启动命令: ```bash export TAOTOKEN_API_KEY=sk-你的Key litellm --config config.yaml --port 4000看到Uvicorn running on http://0.0.0.0:4000就说明代理起来了。
4. 验证工具调用请求与成功返回结构
配置跑起来之后,必须实际发一个带tools的请求,确认返回里真的有tool_calls。这一步不能省,因为很多问题只有真实请求才暴露。
用 curl 打 LiteLLM 的本地端口:
curl http://localhost:4000/v1/chat/completions \ -H "Authorization: Bearer sk-你的TaoTokenKey" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek-chat", "messages": [ {"role": "user", "content": "帮我查一下北京现在的天气"} ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "查询指定城市的天气", "parameters": { "type": "object", "properties": { "city": {"type": "string", "description": "城市名称"} }, "required": ["city"] } } } ], "tool_choice": "auto" }'期望的返回结构里,choices[0].finish_reason应该是tool_calls,choices[0].message.tool_calls是一个数组,里面能看到function.name为get_weather,function.arguments是{"city": "北京"}这样的 JSON 字符串。如果返回的是finish_reason: "stop"且content里是一段解释文字,说明适配没生效,回到第 5 节排查。
Python 客户端验证更贴近 Agent 框架的真实调用方式:
from openai import OpenAI client = OpenAI( base_url="http://localhost:4000/v1", api_key="sk-你的TaoTokenKey", ) response = client.chat.completions.create( model="deepseek-chat", messages=[{"role": "user", "content": "帮我查一下上海现在的天气"}], tools=[{ "type": "function", "function": { "name": "get_weather", "description": "查询指定城市的天气", "parameters": { "type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"], }, }, }], ) choice = response.choices[0] print("finish_reason:", choice.finish_reason) if choice.message.tool_calls: for tc in choice.message.tool_calls: print("tool name:", tc.function.name) print("arguments:", tc.function.arguments)跑通的话输出类似:
finish_reason: tool_calls tool name: get_weather arguments: {"city": "上海"}拿到这个结果,说明整条链路通了。接下来 Agent 框架会执行本地get_weather函数,把结果以role: "tool"的消息追加到对话里,再发一次请求。第二次请求经过async_pre_call_hook时,role: "tool"的消息会被转换成role: "user"加Function: get_weather, Result: ...的文本,Deepseek 就能基于这个结果继续生成最终回答。这个转换逻辑在tool_handler.py里已经写好了,不用额外配置。
如果你用的是 LangChain,把ChatOpenAI的base_url指向http://localhost:4000/v1,api_key填 TaoToken 的 Key,model填deepseek-chat,其余代码不用动。Cline、Continue 这类工具同理,它们底层都是 OpenAI 兼容客户端,改 Base URL 和 Key 就行。
5. 常见报错排查:401、local proxy failed、choices 解析失败
实际跑的时候大概率会撞上几个典型报错,这里按我遇到过的频率排一下。
401 Authentication Error。这个最常见,原因通常是 Key 没传对或者环境变量没生效。检查三处:config.yaml里写的是os.environ/TAOTOKEN_API_KEY,启动 LiteLLM 的 shell 里有没有export TAOTOKEN_API_KEY=sk-...,以及 curl 请求头里的Authorization是不是Bearer sk-...格式。注意 LiteLLM 代理层和 TaoToken 层是两次鉴权,客户端打 LiteLLM 用的 Key 和 LiteLLM 打 TaoToken 用的 Key 可以不同,但都得有效。如果 LiteLLM 配了master_key,客户端请求头要带 master key。
local proxy failed / Connection refused。这个报错说明 LiteLLM 代理没起来或者端口不对。先确认litellm --config config.yaml --port 4000进程还在,curl http://localhost:4000/health能返回。如果用了 Docker 或者远程机器,注意localhost要换成实际 IP。还有一种情况是api_base写成了https://taotoken.net/api少了/v1,LiteLLM 拼出来的 URL 是/api/chat/completions,TaoToken 那边不认这个路径,表现就是连接被拒或者 404。
reading choices / KeyError choices。这个报错出现在async_post_call_success_hook里,说明response.choices是空的或者结构不对。常见原因是 TaoToken 返回了错误响应,比如模型 ID 写错、额度不足、请求体格式有问题,但 LiteLLM 还是走了 success hook。排查方法是在 hook 开头加一行verbose_proxy_logger.debug(f"raw response: {response}"),看实际返回内容。如果response里是error字段,那就是上游的问题,跟适配逻辑无关。另外content为None时content.split会抛AttributeError,可以在解析前加个判空。
OAuth / token 过期类报错。TaoToken 的 Key 一般不会过期,但如果你在控制台手动删了或者重置了,旧 Key 就会 401。重新生成一个,更新环境变量,重启 LiteLLM 即可。注意 LiteLLM 有缓存,改完环境变量一定要重启进程,光export不重启不生效。
工具调用返回了但 arguments 解析失败。这种情况通常是模型输出的 JSON 代码块格式不标准,比如用了单引号、多了尾逗号、或者代码块标记不是```json。prompt_template.j2里的示例要写清楚,让模型照着格式输出。如果还是不稳定,可以在async_post_call_success_hook里加一层容错,用正则提取{...}再json.loads,失败就记日志跳过。
排查的时候有个通用技巧:把 LiteLLM 的日志级别调到 DEBUG,litellm_settings里加set_verbose: true,能看到每个请求的完整入参和出参,定位问题快很多。
6. 长期跑 Agent 的接入建议与资源入口
验证通过之后,如果你打算把这套配置长期用在生产或者日常开发里,有几个点值得注意。
第一,function_cache用的是本地内存缓存,TTL 180 秒。多实例部署时缓存不共享,tool_call_id在实例 A 写入、实例 B 读不到,会导致role: "tool"的消息转换失败。解决办法是换成 Redis 缓存,LiteLLM 的Cache支持type="redis",配一下redis_host和redis_port就行。单机开发不用管这个。
第二,工具描述模板里的示例 JSON 要和你实际工具的参数结构对齐。模板里写的是send_email,你实际用get_weather,模型可能会被示例带偏。最稳妥的做法是模板里只保留通用的格式说明,具体参数结构由tools数组里的parameters字段提供,模型能读懂 JSON Schema。
第三,Deepseek 的 reasoner 模型在工具调用场景下响应会慢一些,因为要先推理再输出。如果 Agent 对延迟敏感,优先用deepseek-chat。TaoToken 的统一通道对这两个模型都支持,切换只需要改model字段。
第四,Agent 框架那边如果开了parallel_tool_calls,当前适配逻辑是逐个解析代码块,能处理多个工具调用,但tool_call_id的生成依赖 LiteLLM 内部机制,建议实测一下多工具并发返回的顺序是否符合预期。
资源入口按用途分一下:需要创建或管理 Key 去控制台,地址是 https://taotoken.net/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api_keys&utm_campaign=rewrite ;想先不写代码直接试模型对话效果,用 https://taotoken.net/model-chat?utm_source=taotoken_aicg_blog_end&utm_content=model_chat&utm_campaign=rewrite ;接入过程中遇到协议细节问题查文档 https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite ;如果你打算长期跑编码类 Agent、需要更稳定的额度和并发,看 Coding Plan https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding_plan&utm_campaign=rewrite 。官网首页 https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= 有完整的模型列表和通道说明。
最后说个实际经验:这套适配方案不只对 Deepseek 有效,任何 OpenAI 兼容但工具调用不完整的模型,只要把model_info.use_proxy打开,都能走同一套转换逻辑。换模型的时候改config.yaml里的model和api_base就行,tool_handler.py不用动。TaoToken 的统一 Key 让这个切换成本降到最低,一个 Key 覆盖多个后端,省去为每个模型单独维护凭证的麻烦。