arXiv:2602.01334cs.CV2026-02被引 5

揭示视觉工具强化学习的真实效果:主要减少工具副作用,而非真正掌握工具使用。

What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom

  • 提出分解框架MED,分离模型内在能力与工具带来的影响。
  • 发现性能提升主要来自内在能力进步,工具使用改进有限。
  • 适合关注VLM工具使用机制与可靠性研究的读者。

视觉工具使用强化学习(Vision Tool-Use RL)可为视觉语言模型(VLM)赋予如裁剪缩放等视觉操作能力并带来显著性能提升,但尚不清楚这些提升是源于工具使用能力的增强,还是模型内在能力的演化。本文提出MED(Measure--Explain--Diagnose)框架,从粗到细地解耦内在能力变化与工具诱导效应,将工具引起的性能差异分解为增益与损害两项,并探究其演变机制。在两个具有不同工具先验的VLM上,针对裁剪缩放任务,在六个基准测试中进行检查点级分析,结果表明性能提升主要由内在学习驱动,而工具使用强化学习主要作用于减少工具诱发的损害(如更少调用错误、更弱的工具模式干扰),在纠正内在缺陷方面的工具化修正进展有限。总体而言,在所研究的裁剪缩放设置下,当前视觉工具使用强化学习主要学习的是与工具安全共存,而非真正掌握工具使用。

原文摘要 · Abstract (English)

Vision tool-use reinforcement learning (RL) can equip vision language models with visual operators such as crop-and-zoom and achieves strong performance gains, yet it remains unclear whether these gains are driven by improvements in tool use or evolving intrinsic capabilities. We introduce MED (Measure--Explain--Diagnose), a coarse-to-fine framework that disentangles intrinsic capability changes from tool-induced effects, decomposes the tool-induced performance difference into gain and harm terms, and probes the mechanisms driving their evolution. Across checkpoint-level analyses in the crop-and-zoom setting on two VLMs with different tool priors and six benchmarks, we find that improvements are dominated by intrinsic learning, while tool-use RL mainly reduces tool-induced harm (e.g., fewer call-induced errors and weaker tool schema interference) and yields limited progress in tool-based correction of intrinsic failures. Overall, in the crop-and-zoom setting studied here, current vision tool-use RL learns to coexist safely with tools rather than master them.

视觉语言模型强化学习工具使用能力解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。