1. “Skill”不是函数,而是智能体的“职业资格证”
很多人第一次接触 LLM Agent 开发时,看到文档里写着“定义一个 Skill”,下意识就去写一个 Python 函数:输入 prompt,调用 model,返回结果——完事。我当年也是这么干的,结果在真实项目里跑了三天就崩了:任务超时、参数错乱、返回格式不一致、下游 Agent 解析失败、日志里全是KeyError: 'response'。后来才明白,Skill 不是 API 封装,也不是工具函数,而是一份结构化、可验证、可编排、带契约约束的“职业资格证”。
它要回答四个核心问题:
- 我能做什么?(功能边界,非模糊描述,而是明确的输入/输出 Schema)
- 我承诺做到什么?(可靠性指标:成功率、响应时间、容错等级)
- 我需要什么才能开工?(依赖项:API Key、本地文件路径、环境变量、第三方服务状态)
- 我出错了该怎么收场?(降级策略、重试逻辑、错误分类码、可观测埋点)
这和写一个def get_weather(city)完全不是一回事。后者是程序员视角的“实现”,前者是工程视角的“契约”。你提交的不是代码,是一份可被调度系统自动校验、被编排引擎动态选择、被监控平台实时评估的标准化能力单元。
关键词里反复出现的YAML和Python,恰恰揭示了 Skill 的双轨结构:YAML 是契约声明层(What),Python 是履约执行层(How)。YAML 描述“这个 Skill 能干什么、要什么、怎么算成功”,Python 实现“具体怎么干”。两者缺一不可,且必须严格对齐——YAML 里声明支持location: str,Python 函数签名就必须有location: str;YAML 里说超时阈值是3000ms,代码里就必须有timeout=3的硬约束。
我见过最典型的反例,是某团队把 Skill 写成一个万能call_external_api()函数,所有参数都塞进kwargs,YAML 文件里只有一行description: "调用外部服务"。结果上线后,Agent 编排器根本无法判断这个 Skill 是否适配当前任务——它连输入字段名都不知道,更别说做参数校验或 fallback 选择了。最后只能靠人工硬编码路由规则,彻底丧失了 Agent 的自主性。
所以,“写出好的 Skill”,本质是用工程化思维重新定义“能力”本身:它不是一段能跑通的代码,而是一个具备身份标识、行为契约、生命周期管理的独立能力实体。接下来,我们就从契约设计、执行实现、测试验证、部署集成四个维度,拆解如何真正落地一个“好”的 Skill。
2. YAML 契约:用声明式语言画清能力的“法律边界”
Skill 的 YAML 文件,不是配置文件,而是能力契约(Capability Contract)。它定义的是 Skill 在整个 Agent 系统中的“法律身份”,而非运行时的参数微调。很多开发者把它当成.env的替代品,只填api_key和base_url,这是根本性误读。一个合格的 Skill YAML 必须包含五个强制区块,缺一不可:
2.1 identity:给 Skill 发一张“身份证”
identity: id: "weather-forecast-v2" name: "Weather Forecast Service" version: "2.3.1" author: "infra-team@company.com" license: "MIT"id是全局唯一标识符,必须小写、无空格、无特殊字符(weather_forecast_v2合法,Weather Forecast (v2)非法)。它是 Agent 调度器索引 Skill 的主键,也是日志追踪、指标聚合的根 ID。name是人类可读名称,用于 UI 展示和调试日志,允许空格和符号。version采用语义化版本(SemVer),主版本号(MAJOR)变更意味着输入/输出 Schema 不兼容。Agent 编排器会拒绝将v1.x的 Skill 绑定到要求v2.x的工作流中。author和license不是形式主义——当 Skill 出现数据泄露或合规风险时,这是追溯责任链的唯一依据。
提示:不要用
uuid生成id。id必须具备业务含义,如gis-spatial-analysis、cv2-object-detection。UUID 会导致运维完全无法理解日志中的skill_id: 550e8400-e29b-41d4-a716-446655440000到底对应哪个能力。
2.2 interface:定义“我能接什么活”的精确接口
interface: inputs: - name: "location" type: "string" required: true description: "城市名称,支持中文或英文,如 'Beijing' 或 '北京'" examples: ["Shanghai", "深圳"] - name: "days" type: "integer" required: false default: 7 min: 1 max: 14 description: "预报天数,范围 1-14" outputs: - name: "forecast" type: "array" items: type: "object" properties: date: { type: "string", format: "date" } temperature_high: { type: "number" } temperature_low: { type: "number" } condition: { type: "string" } description: "未来 N 天天气预报数组" errors: - code: "LOCATION_NOT_FOUND" description: "城市名称未识别,请检查拼写或使用标准地名" - code: "API_RATE_LIMIT_EXCEEDED" description: "服务调用频次超限,请稍后重试"这才是真正的契约核心。它强制规定:
- 输入字段的类型、必选性、默认值、取值范围(
min/max对integer,pattern对string); - 输出结构的嵌套层级、每个字段的类型与格式(
format: date表示 ISO 8601 格式); - 所有可能的错误码及其语义,而非笼统的
500 Internal Error。
我踩过的最大坑,是忽略examples字段。某次上线新 Skill,YAML 中inputs只写了type: string,没给examples。Agent 编排器在生成测试用例时,随机生成了"location": "a"这样的单字符输入,结果触发了上游 API 的异常分支,导致整个测试流水线失败。加上examples后,测试框架能自动生成符合业务场景的典型值,大幅降低误报率。
2.3 execution:声明“我需要什么才能开工”
execution: runtime: "python3.10" dependencies: - name: "requests" version: ">=2.28.0,<2.30.0" - name: "pydantic" version: ">=2.5.0" environment: - name: "WEATHER_API_KEY" required: true description: "Weather API 的密钥,需在运行环境预置" - name: "WEATHER_BASE_URL" required: false default: "https://api.weather.com/v3" timeout_ms: 3000 max_retries: 2runtime明确指定 Python 版本,禁止使用latest或3.x这类模糊标识。不同 Python 小版本间存在 ABI 兼容性差异(如3.10与3.11的asyncio行为变化),必须锁定。dependencies使用 PEP 440 兼容的版本约束,>=2.28.0,<2.30.0比~=2.28.0更安全,避免意外升级到破坏性版本。environment列出所有运行时依赖的环境变量,required: true的变量若缺失,Skill 启动即失败,不会进入执行阶段——这是最廉价的防御性检查。timeout_ms和max_retries是 SLA 的硬性承诺,必须与实际代码中的requests.timeout和重试逻辑完全一致。YAML 里写3000,代码里却用timeout=5,就是欺诈性契约。
2.4 metadata:标注“我在系统里的角色定位”
metadata: category: "data-retrieval" tags: ["weather", "geolocation", "forecast"] capabilities: - "real-time-query" - "batch-processing" reliability: success_rate_target: 0.995 p95_latency_ms: 2500category是 Agent 调度器进行粗粒度路由的依据(如>documentation: overview: "提供全球主要城市的 7 天天气预报,数据源为 Weather.com API。支持中文城市名解析。" usage_examples: - input: { location: "Beijing", days: 3 } output: { forecast: [...] } - input: { location: "上海" } output: { forecast: [...] } security_notes: "不处理用户个人位置信息,所有 location 参数仅用于 API 查询,不记录、不存储。"这不是可选的“文档补充”,而是契约的组成部分。
usage_examples是自动化测试的黄金数据源,security_notes是合规审计的直接证据。没有security_notes的 Skill,在金融、医疗等强监管行业会被直接拒入生产环境。3. Python 执行层:契约落地的“履约引擎”
YAML 契约再完美,最终还是要靠 Python 代码来兑现。但这里的 Python,不是自由发挥的脚本,而是严格遵循契约的履约引擎。它必须完成三件事:精准解析输入、可靠执行逻辑、规范构造输出。任何偏离契约的行为,都是对整个 Agent 系统稳定性的破坏。
3.1 输入解析:契约即真理,绝不妥协
from pydantic import BaseModel, Field, ValidationError from typing import List, Dict, Optional import logging logger = logging.getLogger(__name__) class WeatherInput(BaseModel): location: str = Field(..., min_length=2, max_length=50) days: int = Field(7, ge=1, le=14) class WeatherOutput(BaseModel): forecast: List[Dict[str, any]] # 实际应定义更细的 Pydantic Model def execute(input_data: dict) -> dict: try: # 1. 严格按 YAML interface 定义的 Schema 解析输入 parsed_input = WeatherInput(**input_data) except ValidationError as e: # 2. 错误必须映射到 YAML errors 中定义的 code logger.error(f"Input validation failed: {e}") raise ValueError("LOCATION_NOT_FOUND") from e # 3. 检查运行时依赖(environment 中声明的) api_key = os.getenv("WEATHER_API_KEY") if not api_key: raise ValueError("API_KEY_MISSING") # ... 执行实际逻辑关键点:
- 必须使用 Pydantic v2+ 的
BaseModel进行输入校验,而非if isinstance(input, dict)这类弱类型检查。Field(..., min_length=2)直接对应 YAML 中min_length的约束。 - 错误抛出必须是 YAML
errors中预定义的code字符串,而非Exception类名。Agent 运行时会捕获该字符串,匹配到 YAML 中的code,从而返回标准化的错误响应。 - 环境变量检查必须在业务逻辑前完成。
os.getenv("WEATHER_API_KEY")若为None,立即raise ValueError("API_KEY_MISSING"),而不是等到 API 调用时才报错——后者会让错误归因变得困难。
3.2 执行逻辑:在契约框架内“安全驾驶”
import requests import time from tenacity import retry, stop_after_attempt, wait_exponential @retry( stop=stop_after_attempt(3), # max_retries + 1 wait=wait_exponential(multiplier=1, min=1, max=10), reraise=True ) def _call_weather_api(location: str, days: int, api_key: str) -> dict: start_time = time.time() try: response = requests.get( f"{os.getenv('WEATHER_BASE_URL', 'https://api.weather.com/v3')}/forecast", params={"location": location, "days": days}, headers={"Authorization": f"Bearer {api_key}"}, timeout=3.0 # 严格匹配 YAML timeout_ms / 1000 ) response.raise_for_status() return response.json() except requests.Timeout: logger.warning(f"API timeout for {location}") raise requests.Timeout("API call timed out") except requests.RequestException as e: logger.error(f"API request failed for {location}: {e}") raise e finally: duration_ms = (time.time() - start_time) * 1000 # 上报 P95 延迟指标 if duration_ms > 2500: # p95_latency_ms logger.warning(f"Latency exceeded P95: {duration_ms:.0f}ms") def execute(input_data: dict) -> dict: # ... 输入解析与环境检查(见上节) try: raw_data = _call_weather_api( parsed_input.location, parsed_input.days, api_key ) # 4. 输出构造:必须严格匹配 YAML interface.outputs 定义的结构 forecast_data = _parse_raw_response(raw_data) # 自定义解析函数 return {"forecast": forecast_data} except requests.Timeout: raise ValueError("API_TIMEOUT") except requests.HTTPError as e: if e.response.status_code == 404: raise ValueError("LOCATION_NOT_FOUND") elif e.response.status_code == 429: raise ValueError("API_RATE_LIMIT_EXCEEDED") else: raise ValueError("API_UNEXPECTED_ERROR")- 重试逻辑必须与 YAML
max_retries严格对齐。@retry(stop=stop_after_attempt(3))对应max_retries: 2(首次尝试 + 2 次重试 = 3 次总尝试)。 timeout=3.0必须等于YAML timeout_ms / 1000。这是 SLA 的底线,不能有任何浮动。- 错误分类必须映射到 YAML errors。HTTP 404 →
"LOCATION_NOT_FOUND",429 →"API_RATE_LIMIT_EXCEEDED",确保 Agent 能根据错误码执行预设的 fallback 策略(如 404 时尝试拼音纠错,429 时降级到缓存)。 - 性能监控必须嵌入执行路径。
duration_ms计算和 P95 超标日志,是reliability.p95_latency_ms的实证依据。
3.3 输出构造:契约即宪法,结构不容篡改
def _parse_raw_response(raw: dict) -> List[Dict]: """ 将原始 API 响应转换为 YAML interface.outputs 定义的 forecast 数组 必须保证: - 每个元素是 dict,且包含 date, temperature_high, temperature_low, condition 四个 key - date 是 YYYY-MM-DD 格式字符串 - temperature_* 是 number 类型 - condition 是 string """ forecast_list = [] for day in raw.get("forecast", []): try: # 强制类型转换,确保契约履行 item = { "date": str(day.get("date", ""))[:10], # 截断为 date 格式 "temperature_high": float(day.get("temp_max", 0)), "temperature_low": float(day.get("temp_min", 0)), "condition": str(day.get("condition", "unknown")) } forecast_list.append(item) except (ValueError, TypeError) as e: logger.warning(f"Failed to parse forecast item: {e}, skipping...") continue # 跳过单条异常数据,不中断整个 Skill return forecast_list- 必须进行显式类型转换。API 返回的
temp_max可能是字符串"25",必须float()转为数字,否则下游 Agent 解析 JSON 时会因类型不符而失败。 - 必须处理字段缺失。
day.get("date", "")提供默认值,避免KeyError;[:10]确保date符合format: date要求。 - 单条数据异常必须
continue,而非raise。Skill 的职责是交付尽可能多的有效数据,而非因一条脏数据而整体失败。这是容错性的基本体现。
4. 测试验证:用契约驱动的“法庭审判”
写完 YAML 和 Python,绝不能直接扔进 Agent 系统。一个未经充分验证的 Skill,就像一张未经公证的合同——纸面上很美,执行时全是纠纷。测试不是 QA 环节,而是契约履行的法庭审判,必须覆盖三个维度:契约合规性、功能正确性、鲁棒性。
4.1 契约合规性测试:YAML 的“语法与语义审查”
# test_contract_compliance.py import yaml import jsonschema from jsonschema import validate import pytest def test_yaml_schema(): """验证 YAML 文件是否符合 Skill 契约 Schema""" with open("weather-skill.yaml") as f: data = yaml.safe_load(f) # 加载预定义的 Skill 契约 JSON Schema with open("skill-contract-schema.json") as f: schema = json.load(f) # 执行 JSON Schema 验证 validate(instance=data, schema=schema) assert data["identity"]["id"] == "weather-forecast-v2" assert data["execution"]["timeout_ms"] == 3000 def test_interface_consistency(): """验证 YAML interface 与 Python 类型注解的一致性""" from weather_skill import WeatherInput, WeatherOutput yaml_data = load_yaml("weather-skill.yaml") # 检查 inputs 字段名与 WeatherInput 的字段名是否完全一致 yaml_inputs = [i["name"] for i in yaml_data["interface"]["inputs"]] pydantic_fields = list(WeatherInput.model_fields.keys()) assert set(yaml_inputs) == set(pydantic_fields) # 检查 outputs 结构是否匹配 assert yaml_data["interface"]["outputs"][0]["name"] == "forecast" assert WeatherOutput.model_fields["forecast"].annotation == List[Dict]- JSON Schema 验证:确保 YAML 文件结构合法,
identity.id存在、interface.inputs是数组等基础语法。 - YAML-Python 字段一致性检查:自动比对
YAML interface.inputs.name和Pydantic Model field names,防止手误导致契约与实现脱节。这是最常发生的错误——改了 YAML 忘改 Python,或反之。
4.2 功能正确性测试:用 YAML examples 驱动的“黄金路径”
# test_functional.py import pytest from weather_skill import execute @pytest.mark.parametrize("input_data,expected_output", [ # 直接从 YAML documentation.usage_examples 提取 ({"location": "Beijing", "days": 3}, {"forecast": [...]}) ]) def test_golden_path(input_data, expected_output): """执行 YAML 中声明的黄金用例""" result = execute(input_data) assert "forecast" in result assert len(result["forecast"]) == 3 # 检查 forecast 中每个元素的结构 for item in result["forecast"]: assert "date" in item and isinstance(item["date"], str) assert "temperature_high" in item and isinstance(item["temperature_high"], (int, float)) def test_default_days(): """测试 YAML 中声明的 default 值是否生效""" result = execute({"location": "Shanghai"}) assert len(result["forecast"]) == 7 # default: 7- 参数化测试(
@pytest.mark.parametrize):直接将 YAMLusage_examples中的input/output转为测试用例,确保文档即测试,测试即文档。 - default 值验证:专门测试
days字段不传时是否返回 7 天数据,这是契约中required: false, default: 7的直接验证。
4.3 鲁棒性测试:模拟“法庭上的刁难”
# test_robustness.py import pytest from weather_skill import execute def test_invalid_location(): """测试 YAML errors 中 LOCATION_NOT_FOUND 的触发""" with pytest.raises(ValueError) as exc_info: execute({"location": "NonExistentCity123"}) assert str(exc_info.value) == "LOCATION_NOT_FOUND" def test_timeout_simulation(): """模拟 API 超时,验证 timeout_ms 和重试逻辑""" # monkey patch requests.get to raise Timeout import requests original_get = requests.get def mock_get(*args, **kwargs): raise requests.Timeout("Simulated timeout") requests.get = mock_get try: with pytest.raises(requests.Timeout): execute({"location": "Beijing"}) finally: requests.get = original_get def test_environment_missing(): """测试 environment.required: true 的缺失场景""" import os original_key = os.environ.get("WEATHER_API_KEY") if "WEATHER_API_KEY" in os.environ: del os.environ["WEATHER_API_KEY"] try: with pytest.raises(ValueError) as exc_info: execute({"location": "Beijing"}) assert str(exc_info.value) == "API_KEY_MISSING" finally: if original_key: os.environ["WEATHER_API_KEY"] = original_key- 错误码精准触发:
test_invalid_location确保非法输入必然抛出"LOCATION_NOT_FOUND",而非其他模糊错误。 - 超时与重试验证:
test_timeout_simulation通过 Monkey Patch 模拟网络超时,验证max_retries是否生效、timeout_ms是否被遵守。 - 环境依赖缺失测试:
test_environment_missing删除WEATHER_API_KEY,验证 Skill 在启动阶段就失败,而非在 API 调用时才崩溃。
4.4 集成测试:在 Agent 环境中“走一遍全流程”
# test_integration.py from agent_core.scheduler import SkillScheduler from agent_core.executor import SkillExecutor def test_skill_in_scheduler(): """验证 Skill 能被 Agent 调度器正确加载和识别""" scheduler = SkillScheduler() scheduler.load_skill_from_yaml("weather-skill.yaml") # 检查调度器是否能根据 category 和 tags 找到 Skill candidates = scheduler.find_skills( category="data-retrieval", tags=["weather"] ) assert len(candidates) == 1 assert candidates[0].id == "weather-forecast-v2" def test_end_to_end_execution(): """在模拟 Agent 运行时环境中执行完整流程""" executor = SkillExecutor() # 模拟 Agent 传入的标准化输入 agent_input = { "skill_id": "weather-forecast-v2", "parameters": {"location": "Beijing", "days": 3} } result = executor.execute(agent_input) assert result["status"] == "success" assert "forecast" in result["output"] assert result["latency_ms"] < 3000 # 验证 timeout_ms 约束- 调度器集成:验证 Skill 能被
SkillScheduler正确加载,并能根据category/tags被检索到,这是 Agent 自主决策的基础。 - 端到端执行:在
SkillExecutor(Agent 的实际执行引擎)中运行,验证从 Agent 输入到 Skill 输出的完整链路,包括latency_ms等运行时指标的采集。
5. 部署与可观测:让 Skill 在生产中“透明可信”
一个 Skill 上线后,就不再是孤立的代码,而是 Agent 系统的有机组成部分。它的健康状况、性能表现、错误模式,必须对运维、开发、产品团队完全透明。这需要一套契约驱动的可观测体系,其核心是:所有监控指标、日志字段、追踪 Span,都必须能回溯到 YAML 契约中的声明。
5.1 指标(Metrics):用 YAML reliability 驱动的监控看板
# metrics.py from prometheus_client import Counter, Histogram, Gauge # 基于 YAML metadata.category 和 identity.id 构建指标 SKILL_SUCCESS_TOTAL = Counter( 'skill_success_total', 'Total number of successful skill executions', ['skill_id', 'category'] ) SKILL_FAILURE_TOTAL = Counter( 'skill_failure_total', 'Total number of failed skill executions', ['skill_id', 'category', 'error_code'] # error_code 来自 YAML errors.code ) SKILL_LATENCY_SECONDS = Histogram( 'skill_latency_seconds', 'Latency of skill execution', ['skill_id', 'category'], buckets=[0.001, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0] ) def record_execution(skill_id: str, category: str, success: bool, error_code: str = None, latency_ms: float = 0.0): if success: SKILL_SUCCESS_TOTAL.labels(skill_id=skill_id, category=category).inc() else: SKILL_FAILURE_TOTAL.labels( skill_id=skill_id, category=category, error_code=error_code or "UNKNOWN" ).inc() # 转换为秒并记录 Histogram SKILL_LATENCY_SECONDS.labels(skill_id=skill_id, category=category).observe(latency_ms / 1000.0)- 指标标签(Labels)必须包含
skill_id和category,这是多维分析的基础。你可以轻松查询:“weather-forecast-v2在过去 1 小时内的成功率是多少?”或“># logger.py import logging import json from datetime import datetime class SkillJsonFormatter(logging.Formatter): def format(self, record): log_entry = { "timestamp": datetime.utcnow().isoformat(), "level": record.levelname, "skill_id": getattr(record, 'skill_id', 'unknown'), "category": getattr(record, 'category', 'unknown'), "input": getattr(record, 'input', {}), "output": getattr(record, 'output', {}), "error_code": getattr(record, 'error_code', None), "latency_ms": getattr(record, 'latency_ms', 0), "message": record.getMessage() } return json.dumps(log_entry) # 在 execute 函数中使用 def execute(input_data: dict) -> dict: start_time = time.time() try: # ... 执行逻辑 result = {"forecast": forecast_data} # 记录成功日志,包含 input 和 output logger.info("Skill executed successfully", extra={ "skill_id": "weather-forecast-v2", "category": "data-retrieval", "input": input_data, "output": result, "latency_ms": (time.time() - start_time) * 1000 }) return result except ValueError as e: # 记录错误日志,包含 error_code logger.error("Skill execution failed", extra={ "skill_id": "weather-forecast-v2", "category": "data-retrieval", "input": input_data, "error_code": str(e), "latency_ms": (time.time() - start_time) * 1000 }) raise- 日志必须是 JSON 格式,且字段名与 YAML 契约对齐(
skill_id,input,error_code)。 input和output字段必须记录,这是调试的核心线索。当 Agent 报错时,运维人员可以直接在日志中看到“当时传了什么输入,得到了什么输出(或错误)”。extra参数传递上下文,避免在日志消息字符串中拼接,保证结构化。
5.3 追踪(Tracing):用 YAML identity 构建的分布式调用链
# tracing.py from opentelemetry import trace from opentelemetry.trace import Status, StatusCode tracer = trace.get_tracer(__name__) def execute_with_trace(input_data: dict) -> dict: with tracer.start_as_current_span("weather-forecast-v2.execute") as span: # 设置 Span 属性,全部来自 YAML span.set_attribute("skill.id", "weather-forecast-v2") span.set_attribute("skill.version", "2.3.1") span.set_attribute("skill.category", "data-retrieval") span.set_attribute("skill.input.location", input_data.get("location", "")) span.set_attribute("skill.input.days", input_data.get("days", 7)) try: result = _execute_core_logic(input_data) span.set_attribute("skill.output.forecast_count", len(result["forecast"])) span.set_status(Status(StatusCode.OK)) return result except ValueError as e: span.set_attribute("skill.error.code", str(e)) span.set_status(Status(StatusCode.ERROR)) raise- Span 名称必须是
skill_id.execute,如weather-forecast-v2.execute,这是 Jaeger/Grafana Tempo 中搜索调用链的唯一标识。 - Span 属性(Attributes)必须包含 YAML identity 和 interface 的关键字段,如
skill.id,skill.version,skill.input.location。这样就能在分布式追踪系统中,按 Skill ID、版本、甚至特定输入参数(如location=Beijing)进行筛选和分析。
5.4 健康检查(Health Check):YAML execution 的实时心跳
# health.py from fastapi import APIRouter, Response import os router = APIRouter() @router.get("/health/skill/{skill_id}") def skill_health_check(skill_id: str): """ 健康检查端点,验证 Skill 的 runtime 和 dependencies 是否就绪 """ # 1. 检查 Python runtime 版本 import sys if not sys.version.startswith("3.10"): return Response(status_code=503, content="Python version mismatch") # 2. 检查 dependencies(通过 import) try: import requests import pydantic except ImportError as e: return Response(status_code=503, content=f"Missing dependency: {e}") # 3. 检查 environment variables(来自 YAML execution.environment) required_envs = ["WEATHER_API_KEY"] for env in required_envs: if not os.getenv(env): return Response(status_code=503, content=f"Missing environment: {env}") # 4. 执行轻量级自检(如 ping API) try: # 不调用完整逻辑,只做最小可行性验证 import requests response = requests.get( f"{os.getenv('WEATHER_BASE_URL', 'https://api.weather.com/v3')}/health", timeout=1.0 ) response.raise_for_status() except Exception as e: return Response(status_code=503, content=f"External service unhealthy: {e}") return {"status": "ok", "skill_id": skill_id, "timestamp": time.time()}- 健康检查必须覆盖 YAML
execution的所有声明:runtime、dependencies、environment、timeout_ms。 - 必须包含对外部服务的轻量级探测(如
/health端点),而非仅检查本地依赖。一个 Skill 的健康,取决于其整个依赖链的健康。 - HTTP 状态码必须准确:
503 Service Unavailable表示 Skill 不可用,200 OK表示一切正常。Kubernetes 的 Liveness Probe 会据此决定是否重启 Pod。
6. 工程实践:从“写 Skill”到“运营 Skill”的认知跃迁
写一个能跑通的 Skill,可能只需要一小时;但写一个能在生产环境稳定运行、被多个 Agent 共享、经得起流量洪峰和异常冲击的 Skill,需要一整套工程方法论。这背后,是开发者角色从“代码实现者”到“能力运营者”的跃迁。
6.1 版本管理:Semantic Versioning 是契约演进的宪法
Skill 的版本号 `2.
- 日志必须是 JSON 格式,且字段名与 YAML 契约对齐(
- 必须使用 Pydantic v2+ 的