arXiv:2608.08557cs.CLcs.CV2026-08被引 1

构建可解释的视觉工具使用数据集,让模型学会用工具获取证据而非模仿动作。

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

论文配图:OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories
图 1 · 摘自论文原文
  • 通过因果有效性筛选工具使用轨迹,确保动作真正影响答案。
  • 生成42K条高质量轨迹,覆盖5个视觉推理领域,提升模型表现。
  • 适合研究多模态智能体、视觉推理与可控生成的开发者与研究者。

视觉工具使用已成为多模态智能体主动获取图像编码之外证据的关键能力。现有方法依赖教师生成并筛选正确答案的轨迹,隐含假设所有成功示范都提供有效监督。我们指出该假设存在问题:强教师常无需工具即可得出正确答案,模仿此类轨迹会使学生误以为工具调用只是正确答案的伴随行为,而非证据获取的必要步骤。为此,提出OpenVisTool——一个开放框架,用于构建具有指导意义的视觉工具使用轨迹。核心思想是:仅保留答案正确(结果有效性)且工具观察对答案有因果贡献(因果效用)的轨迹。框架分三阶段:难度筛选(选择非工具不可解的问题)、领域特定轨迹生成(诱导连贯工具使用)、监督验证(联合测试双条件)。不同于鼓励模仿工具调用,新监督教会模型何时何地获取视觉证据。基于此,构建了覆盖五个视觉推理领域的OpenVisTool-42K数据集及对应OpenVisTool-Bench基准。在四个模型规模(4B-27B)上微调均显著提升视觉工具使用性能,并在两个分布外基准上取得增益;大模型接近领先闭源系统。结果表明,有效的视觉工具使用源于因果驱动的监督,而非工具调用模式。

原文摘要 · Abstract (English)

Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision. We argue this assumption is flawed: a strong teacher often reaches the correct answer without needing its tool calls, and imitating such trajectories teaches a student that tool calls accompany correct answers, not that tool observations ground them. We present OpenVisTool, an open framework for constructing instructive visual tool-use trajectories that provide effective supervision for tool learning. The key insight is that a trajectory should be retained only if its answer is correct (outcome validity) and its tool observations causally contribute to that answer (causal utility). The framework operates in three stages: difficulty screening to select queries that are not reliably answerable without tools, domain-specific trajectory synthesis to elicit coherent tool-use trajectories, and supervision verification to jointly test both conditions. Rather than encouraging models to imitate tool calls, the resulting supervision teaches when and how visual evidence should be acquired. Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains. Across four backbones (4B-27B), fine-tuning on OpenVisTool-42K consistently improves visual tool-use performance and yields gains on two out-of-distribution benchmarks; the larger models approach leading closed-source systems. The evidence suggests that effective visual tool use is learned from causally grounded supervision rather than tool-calling patterns.

视觉推理工具使用数据构建因果学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。