构建可诊断的工具生成基准,评估语言模型从需求自动生成可用工具的能力。
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
- 基于抽象任务需求动态生成工具接口与逻辑,不依赖预设规范。
- 顶尖模型在一次性生成中仍难以保证接口精确性和逻辑可执行性。
- 揭示初始微小缺陷会级联放大,导致下游性能严重下降,适合工具链研究者使用。
自进化语言代理研究快速发展,其核心能力在于根据任务需求创建、适配和维护工具。然而现有基准多依赖预定义规范,限制了可扩展性,阻碍真正自主演化。尽管近期研究尝试动态生成工具,但主要关注下游性能,导致评估呈现“黑箱”特性,难以定位失败原因。为此,我们提出Tool-Genesis,一个诊断性基准,用于量化代理在接口合规性、功能正确性和下游实用性等多个维度的表现。该基准评估代理是否仅凭抽象需求(无预设规范)即可构建相关工具并用于解决真实问题。关键发现:即使最先进的模型在单次生成场景下也难以产出精确的工具接口或可执行逻辑。这些初始微小缺陷在流水线中被放大,导致下游指标急剧下降。我们期望Tool-Genesis能引导未来研究,推动模型训练与调控,以合成持久、通用的工具,更好地应对现实挑战。
原文摘要 · Abstract (English)
Research on self-evolving language agents has accelerated, drawing increasing attention to their ability to create, adapt, and maintain tools from task requirements. However, existing benchmarks predominantly rely on predefined specifications, which limits scalability and hinders truly autonomous evolution. While recent studies attempt to dynamically generate tools, they primarily emphasize downstream performance, resulting in a "black-box" evaluation that makes it difficult to attribute failures to specific causes. To address this, we propose Tool-Genesis, a diagnostic benchmark designed to quantify agent capabilities across multiple dimensions, including interface compliance, functional correctness, and downstream utility. Tool-Genesis evaluates whether agents can construct task-relevant tools solely from abstract requirements (without preset specifications) and use them to solve realistic problems. Crucially, we find that even state-of-the-art models struggle to produce precise tool interfaces or executable logic in a one-shot setting. These minor initial flaws are amplified through the pipeline, leading to a sharp degradation in downstream metrics. We hope Tool-Genesis will guide future research toward training and steering models to synthesize persistent, general-purpose tools that better address real-world challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。