工具增强型多模态智能体未必真会用工具,多数能力提升来自模仿调用模式。
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains

- 对比带工具和无工具版本,验证工具是否真正带来能力提升
- 93%~96%的题目在非工具设置下也能解决,工具贡献有限
- 工具调用行为更像模式模仿,而非实际利用工具信息
工具增强的多模态智能体在基准测试中表现优异,常被视作学会使用工具的证据。我们认为这种解释可能过早:仅凭工具调用记录,并不能说明工具提供了关键答案信息。本文系统研究了两种代表性‘用图像思考’的智能体——Thyme 和 DeepEyesV2——在真实世界理解、OCR、图表理解及数学推理任务中的表现。每个智能体均与无工具版本及从相同数据池训练但无工具调用轨迹的纯文本推理器进行对比。结果显示,工具接入并未带来稳定的总体提升,未显著降低生成词元成本,且仅有少量问题仅由工具解决:DeepEyesV2 的 93% 工具解题项、Thyme 的 96% 可被至少一种非工具设置解决。机制消融实验进一步表明,完整工具使用流程并不优于仅工具调用格式或返回执行结果本身。在本研究设定下,分析的智能体似乎更擅长学习工具调用模式,而非真正利用工具带来的能力扩展,提示评估应区分工具可用性与工具是否实际扩大解决能力。
原文摘要 · Abstract (English)
Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interpretation can be premature: a tool-call trace alone does not show whether the tool supplied answer-critical information. We study two representative ``thinking with images'' agents, Thyme and DeepEyesV2, across real-world understanding, OCR, chart understanding, and mathematical reasoning. Each agent is compared with its Tool-Free counterpart and with a Pure-Text Reasoner trained from the same source pool without tool-calling trajectories. Tool access yields little consistent aggregate improvement, does not reliably reduce generated-token cost, and leaves only a small tool-only solved set: 93% of DeepEyesV2's tool-solved problems and 96% of Thyme's are also solved by at least one non-tool setting. Mechanism ablations further show that the full tool-use loop does not consistently outperform either the tool-call format or the returned execution result alone. In the settings we study, the analyzed agents appear to learn tool-calling patterns more reliably than tool-contributed capabilities, suggesting that evaluation should distinguish tool availability from whether tools actually expand what agents can solve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。