arXiv:2602.01983cs.AI2026-02被引 4

让AI从用工具变成造工具,推理时自动优化能力

Evolving from Tool User to Creator via Training-Free Experience Reuse in Multimodal Reasoning

  • 不训练、不改模型,从历史推理中提炼可复用的工具创建经验
  • 在多领域数学科学任务上性能提升20.86%和23.04%
  • 适合需要持续进化推理能力的智能体系统

现有工具集成推理(TIR)模型通过引入外部工具有效扩展了大模型的问答能力。然而,在现实场景中,开放性问题常超出固定工具的适用范围,且错误工具输出会误导大模型生成结果。此外,工具构建需大量人工投入,限制了其应用。我们发现大模型的推理轨迹蕴含隐式求解能力,提出无需训练的UCT框架,将智能体从工具使用者转变为工具创造者。该方法挖掘推理经验并提炼为可复用资产,实现推理过程中的自适应工具创建与自我更新。我们还设计记忆固化机制,维护工具库,保障经验记忆的高复用性。这一自动化工具构建范式在推理中持续提升工具质量,使整体系统无需额外训练即可演进。大量实验表明,该方法显著提升了TIR模型能力,在跨领域的数学与科学推理基准上分别获得+20.86%↑和+23.04%↑的性能提升,验证了智能体的自演化能力。

原文摘要 · Abstract (English)

Existing Tool-Integrated Reasoning (TIR) models have effectively extended the question-answering capabilities of LLMs by incorporating external tools. However, real-world scenarios present numerous open-ended problems where fixed tools often fail to meet task requirements. Furthermore, the lack of self-optimization mechanisms means that erroneous tool outputs can mislead the LLM's responses. Additionally, the construction of existing tools entails significant manual effort, which consequently constrains their applicability. Recognizing that the reasoning traces of LLMs encapsulate implicit problem-solving capabilities, we propose UCT, a novel training-free framework that transforms agents from tool users to tool creators. This approach harvests reasoning experiences and distills them into reusable assets. This method transforms the agent from a mere tool user into a tool creator, enabling adaptive tool creation and self-updating during the inference process. We also introduce a memory consolidation mechanism to maintain the tool library, ensuring high reusability of retained experiential memory for subsequent reasoning tasks. This novel automated tool construction paradigm continuously improves tool quality during reasoning, allowing the overall agent system to progress without additional training. Extensive experiments demonstrate that our method serves as a novel paradigm for enhancing the capabilities of TIR models. In particular, the significant performance gains achieved +20.86%$\uparrow$ and +23.04%$\uparrow$ on benchmarks across multi-domain mathematical and scientific reasoning tasks validate the self-evolving capability of the agent.

工具生成自演化多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。