Agent 技能与工具体系
本页深入探讨 Agent 的工具使用工程(Tool Use Engineering)与技能体系(Skills Framework)——当 Agent 的基本架构(ReAct 循环、MCP 工具协议)搭建起来后,真正决定 Agent 好不好用的是工具的设计、选择策略和评估方法。这是 Agent 架构模式的工程落地篇。
Agent 的能力 = LLM 推理能力 × 工具能力 × 工具选择准确率。
三个因子缺一不可——推理弱则规划乱,工具少则做不了事,选错工具则白费力气。本页聚焦后两个因子:如何设计好工具、如何让模型选对工具、如何评估整个系统。
- 工具即函数:从 LLM 视角,每个工具就是一个”可调用的函数”——有名字、有参数描述、有返回值格式。模型通过 Function Calling(函数调用)机制选择工具并生成参数。
- 技能 = 工具 + Prompt + 流程:单个工具太原子化,真正解决业务问题需要把多个工具串成流水线,并配套相应的 Prompt 模板和错误处理逻辑——这就是”技能”。
- 评估是瓶颈:Agent 的多步执行使得传统单轮评估(BLEU、ROUGE)完全失效。一个 10 步的 Agent 流程,任何一步出错都会级联放大——需要端到端任务级评估。
Function Calling 是 OpenAI 在 2023 年 6 月推出的标准化接口:LLM 不再输出自然语言文本,而是输出一段符合 JSON Schema 约束的结构化数据,指明”调用哪个函数、传什么参数”。这使 LLM 能与外部 API、数据库、代码执行器精确对接。详见结构化输出。
工具定义:Function Calling 格式
Section titled “工具定义:Function Calling 格式”从工程角度,一个工具由三部分组成:名称(做什么)、参数 Schema(需要什么输入)、描述(什么时候用)。LLM 在推理时读取这些描述,决定是否调用以及如何填充参数。
# 以 OpenAI Function Calling 格式定义工具tools = [ { "type": "function", "function": { "name": "search_web", "description": "Search the web for real-time information. Use this when " "the user asks about current events, facts after your " "training cutoff, or information you are not confident about.", "parameters": { "type": "object", "properties": { "query": { "type": "string", "description": "The search query string" }, "max_results": { "type": "integer", "description": "Maximum number of results to return", "default": 5 } }, "required": ["query"] } } }, { "type": "function", "function": { "name": "sql_query", "description": "Execute a read-only SQL query on the company database. " "Use this to retrieve structured business data like sales, " "users, orders, etc.", "parameters": { "type": "object", "properties": { "query": { "type": "string", "description": "The SQL SELECT query to execute" } }, "required": ["query"] } } }]
# 调用 LLM,让它自主选择工具response = client.chat.completions.create( model="gpt-4o", messages=[ {"role": "user", "content": "查一下上个月销售额最高的 10 个产品"} ], tools=tools, tool_choice="auto" # auto = 模型自己决定是否调用工具)
# 模型输出:自动选择了 sql_query 工具tool_call = response.choices[0].message.tool_calls[0]# tool_call.function.name == "sql_query"# tool_call.function.arguments == '{"query": "SELECT product_name, SUM(amount) ... "}'工具描述是 Prompt 工程的一部分。description 写得好不好直接影响模型选对工具的概率。好的描述需要回答两个问题:这个工具做什么?什么场景下该用它而不是别的工具?
工具选择概率模型
Section titled “工具选择概率模型”从概率视角,工具选择本质是一个分类问题:给定当前对话上下文 和可用的 个工具 ,LLM 输出每个工具的选择概率:
其中 是上下文的隐向量(hidden state), 是工具 的描述嵌入向量, 是温度参数。当工具描述与用户意图在语义空间中对齐时,内积 越大,选择概率越高。
这也解释了为什么工具数量增加时准确率会下降——候选类别 变大,softmax 分母项增多,正确工具的概率被稀释。实践中工具超过 15-20 个时,通常需要引入两级路由(先选工具集,再选具体工具)或 RAG 方式动态检索相关工具。
下表展示了不同查询下各工具的选择概率分布(模拟数据):

工具选择流程
Section titled “工具选择流程”Agent 架构与工具集成
Section titled “Agent 架构与工具集成”一个完整的 Agent 系统不只是 LLM + 工具列表,还包含记忆、规划、错误恢复等模块:
Text2SQL:Agent 技能的典型案例
Section titled “Text2SQL:Agent 技能的典型案例”Text2SQL(自然语言转 SQL)是 Agent 工具使用中最成熟的垂直技能之一——用户说”查一下上季度利润最高的三个产品”,Agent 自动生成 SQL、查询数据库、返回结果。它完整展现了工具使用的工程挑战。
Text2SQL 流水线
Section titled “Text2SQL 流水线”Python 实现:基于 Agent 的 Text2SQL
Section titled “Python 实现:基于 Agent 的 Text2SQL”import json
class Text2SQLAgent: """A minimal Text2SQL agent with schema linking and self-correction."""
def __init__(self, llm_client, db_engine): self.llm = llm_client self.db = db_engine
def get_relevant_schema(self, question: str) -> str: """Step 1: Retrieve relevant table schemas using keyword matching.""" schema_info = self.db.get_full_schema() # all tables/columns # Simple heuristic: include tables whose names appear in the question relevant = [] for table in schema_info: if table["name"].lower() in question.lower(): relevant.append(table) # Fallback: if no match, include all (small DB) or use embedding search if not relevant: relevant = schema_info[:5] # top-5 by frequency return json.dumps(relevant, indent=2)
def generate_sql(self, question: str, schema: str) -> str: """Step 2: Generate SQL using LLM with few-shot examples.""" prompt = f"""You are a SQL expert. Given the database schema and a question,write a safe read-only SQL query.
Database Schema:{schema}
Few-shot Examples:Q: "total sales last month"SQL: SELECT SUM(amount) FROM sales WHERE month = DATE_TRUNC('month', NOW() - INTERVAL '1 month')
Question: "{question}"
Rules:- Only generate SELECT statements- Use table aliases for readability- Return ONLY the SQL, no explanation
SQL:""" response = self.llm.complete(prompt) sql = response.strip().strip('```sql').strip('```').strip() return sql
def validate_sql(self, sql: str) -> tuple[bool, str]: """Step 3: Basic safety validation.""" sql_lower = sql.lower().strip() forbidden = ["drop", "delete", "update", "insert", "alter", "truncate"] for kw in forbidden: if kw in sql_lower: return False, f"Forbidden keyword: {kw}" if not sql_lower.startswith("select"): return False, "Only SELECT queries are allowed" return True, "OK"
def run(self, question: str, max_retries: int = 2): """Full pipeline with self-correction loop.""" schema = self.get_relevant_schema(question)
for attempt in range(max_retries + 1): sql = self.generate_sql(question, schema) is_valid, msg = self.validate_sql(sql) if not is_valid: schema += f"\n\n[Previous SQL was rejected: {msg}. Try again.]" continue
try: rows, columns = self.db.execute(sql) return {"sql": sql, "columns": columns, "rows": rows} except Exception as e: # Self-correction: feed error back to LLM schema += f"\n\n[Previous SQL failed: {e}. Fix the query.]"
return {"error": "Failed after max retries"}
# 使用agent = Text2SQLAgent(llm_client=openai_client, db_engine=pg_engine)result = agent.run("上月销售额最高的 10 个产品")print(result["sql"])# SELECT product_name, SUM(amount) AS total# FROM sales# WHERE date_trunc('month', created_at) = date_trunc('month', now() - interval '1 month')# GROUP BY product_name ORDER BY total DESC LIMIT 10Text2SQL 精度演进
Section titled “Text2SQL 精度演进”Text2SQL 的精度在过去十年持续提升,关键转折在于大模型的引入和 Agent 架构的应用:

Spider 基准(Spider Benchmark)是 Text2SQL 领域最权威的评测集,包含 200+ 数据库、10,000+ 自然语言-SQL 对。Execution Accuracy(EX)衡量执行结果是否正确,Exact Match(EM)衡量 SQL 语法是否完全一致——EX 更贴近实际应用价值。
Agent 评估方法
Section titled “Agent 评估方法”Agent 的多步执行特性使得评估远比单轮 LLM 复杂。业界发展出了多维度评估框架:
1. 任务级指标
Section titled “1. 任务级指标”最核心的指标——整个任务做对了没有:
看似简单,但”成功”的定义因任务而异:对话任务看最终回答是否正确,代码任务看测试是否通过,工具使用任务看最终状态是否达成目标。
2. 过程级指标
Section titled “2. 过程级指标”| 指标 | 含义 | 计算方式 |
|---|---|---|
| Tool Selection Accuracy | 选对工具的比例 | |
| Parameter Accuracy | 参数填对的概率 | |
| Step Efficiency | 实际步数 vs 最优步数 | |
| Error Recovery Rate | 出错后能自我修复的比例 |
3. 端到端延迟
Section titled “3. 端到端延迟”Agent 的多步串行特性导致延迟累积:
其中 是总步数, 是第 步的 LLM 推理延迟, 是工具执行延迟。优化方向包括:减少步数(更好的规划)、并行化(独立步骤同时执行)、缓存(重复查询不重新调用)。
下表用雷达图对比了不同 Agent 配置在各维度上的表现:

4. Agent 评估基准
Section titled “4. Agent 评估基准”业界主流的 Agent 评估基准:
| 基准 | 评测内容 | 特点 |
|---|---|---|
| AgentBench | 多场景 Agent 任务(OS、DB、Web、KG) | 8 类环境,全面评估 |
| WebArena | 网页操作任务(购物、论坛、CMS) | 真实网站环境 |
| SWE-bench | GitHub issue 修复(写代码 + 跑测试) | 软件工程任务,最难 |
| ToolBench / API-Bank | API/工具调用 | 海量真实 API |
| GAIA | 通用 Assistant 任务(需多步推理 + 工具) | 人类均分 92%,模型远低于此 |
SWE-bench 是目前最具挑战性的 Agent 基准之一:给定一个真实 GitHub issue,Agent 需要理解代码库、定位问题、编写修复代码、并通过测试套件。2024 年最好的模型也只能解决约 20-30% 的问题——这直观说明了 Agent 技术距”通用可靠”还有很大差距。详见 LLM 评估。
技能工程:从单工具到技能体系
Section titled “技能工程:从单工具到技能体系”单个工具是原子操作,真正的业务能力需要把多个工具组织成技能(Skill)。技能包含:
- 工具组合:一个技能可能调用多个工具。如”数据分析技能”=
sql_query+python_exec+chart_render。 - Prompt 模板:针对特定任务优化的系统提示。如 Text2SQL 技能有专门的 Schema 感知 Prompt。
- 工作流约束:限定工具调用顺序、参数范围、错误处理策略。
- 输入/输出 Schema:技能对外暴露统一的调用接口。
# 技能定义示例:数据分析技能class DataAnalysisSkill: """组合 SQL 查询 + 数据分析 + 可视化的复合技能。"""
TOOLS = ["sql_query", "python_exec", "chart_render"]
SYSTEM_PROMPT = """You are a data analyst. Follow this workflow:1. Use sql_query to retrieve data from the database2. Use python_exec to clean and analyze the data3. Use chart_render to create visualizations4. Summarize findings for the user
Always start with understanding the schema before writing queries."""
def get_tool_config(self): """返回该技能允许使用的工具子集和约束。""" return { "allowed_tools": self.TOOLS, "max_steps": 15, "constraints": { "sql_query": "read_only: true, max_rows: 10000", "python_exec": "timeout: 30s, memory: 512MB" } }多工具路由:两级选择策略
Section titled “多工具路由:两级选择策略”当工具数量增多时(>20 个),单层 softmax 的选择准确率下降。解决方案是两级路由:
两级路由的选择概率分解为:
先选技能类别( 个候选),再在技能内选具体工具(每个技能 个候选)。这样每个决策点的候选数都控制在合理范围。
工具设计原则
Section titled “工具设计原则”- 高内聚、低耦合:每个工具做一件事,功能边界清晰。不要设计一个”万能查询工具”——模型不知道什么时候该用它。
- 描述要精确:description 要说明三件事——做什么、什么时候用、什么时候不该用。好的描述可以提升 15-20% 的选择准确率。
- 参数要少:必需参数控制在 2-3 个以内。参数越多,模型填错的概率越高。
- 错误信息要可读:工具报错信息要包含修复建议,如
"Column 'price' not found. Available columns: [name, amount, quantity]"——这样 LLM 能自动修正。
- 工具冲突:两个工具功能重叠(如
search_web和google_search),模型随机选择导致不一致——合并或明确区分。 - 参数幻觉:模型生成不存在的参数值(如虚构一个不存在的 column 名)。解决方案:在 description 中列出可选值,或用 约束解码。
- 无限循环:Agent 反复调用同一个工具不收敛。设置
max_steps和重复检测机制。 - 成本失控:每步都调大模型,10 步就是 10 倍成本。对于结构化子任务,考虑用小模型或直接硬编码逻辑。
- SQL 注入防护:Text2SQL 场景中,必须强制只读 + 参数化查询,永远不要让 LLM 生成的 SQL 直接执行 DDL/DML。
- 沙箱隔离:代码执行类工具必须在隔离容器中运行,限制网络访问、文件系统权限和资源用量。
- 权限最小化:MCP 工具只授予完成任务所需的最小权限,避免”给 Agent 一个 root shell”的过度授权。
2025 年趋势
Section titled “2025 年趋势”- Tool Learning 成为独立研究方向:不再将工具使用视为通用 NLP 任务,而是发展出专门的训练方法——Tool-augmented fine-tuning、Tool instruction tuning 等,在工具选择和参数生成上显著优于通用模型。
- MCP 生态爆发:Anthropic 的 MCP(Model Context Protocol)协议在 2025 年获得广泛采用,数千个 MCP Server 覆盖了主流 SaaS 工具(Slack、GitHub、Notion、数据库等),Agent 获取工具的门槛大幅降低。详见 MCP 与工具。
- SWE-bench 突破:2025 年初,多个 Agent 系统在 SWE-bench Verified 上突破 50% 解决率(如 Devin、SWE-Agent),相比 2024 年的 20-30% 有质的飞跃——但仍远低于人类的 92%。
- Agent 训练数据合成:为解决 Agent 训练数据稀缺的问题,合成数据方法兴起——用强模型生成轨迹(trajectory),蒸馏给小模型。OpenAI、Google 都在此方向投入。
- 长程任务规划:超过 20 步的任务,Agent 的规划能力急剧下降——每步 90% 的成功率在 20 步后只剩 12% 的端到端成功率。
- 工具版本管理:API 变更后旧 Agent 可能失效——需要”工具兼容性测试”机制。
- 评估的可重复性:Agent 的随机性(temperature > 0)使评估结果不稳定,同一任务多次运行结果可能不同——业界倾向于多次运行取平均。
- 成本效益:复杂 Agent 任务可能消耗大量 token——需要在精度和成本间权衡,选择合适的模型规模。
Agent 技能体系的未来方向是自主技能学习:Agent 通过试错和反馈自动改进技能的 Prompt 和工具组合,而不需要人工调优——类似于模型层面的”在线学习”。这在 2025 年仍处于早期研究阶段。
- Yao, S. et al. “ReAct: Synergizing Reasoning and Acting in Language Models.” ICLR 2023.
- Schick, T. et al. “Toolformer: Language Models Can Teach Themselves to Use Tools.” NeurIPS 2023.
- Patil, S. et al. “Gorilla: Large Language Model Connected with Massive APIs.” arXiv 2305.15334, 2023.
- Liu, Q. et al. “AgentBench: Evaluating LLMs as Agents.” ICLR 2024.
- Jimenez, C. et al. “SWE-bench: Can Language Models Resolve Real-World GitHub Issues?” ICLR 2024.
- Yu, T. et al. “Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task.” EMNLP 2018.
- OpenAI. “Function Calling and Other API Updates.” June 2023.