arXiv:2503.19263cs.CV2025-03ICCV被引 9

让AI更懂工具使用,提升视觉推理能力

DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning

  • 通过分析工具使用差异生成更有效的推理流程
  • 在多个数据集上达到当前最优性能,泛化能力强
  • 适合需要精准工具调用的视觉推理场景

视觉推理对实现类人级视觉理解至关重要,但依然极具挑战。近期基于大语言模型与工具结合的组合式推理方法展现出优于端到端方法的潜力,但受限于冻结的LLM缺乏工具意识,性能受限。由于训练数据少、工具不完善导致错误、工作流噪声大,直接使用现有方法难以在视觉推理中有效微调。为此,我们提出DWIM:i)差异感知的工作流生成,评估工具使用情况,提取更可行的训练流程;ii)指令掩码微调,引导模型仅复制有效操作,生成更实用的解决方案。实验表明,DWIM在多种视觉推理任务上达到当前最佳表现,并在多个常用数据集上展现强泛化能力。

原文摘要 · Abstract (English)

Visual reasoning (VR), which is crucial in many fields for enabling human-like visual understanding, remains highly challenging. Recently, compositional visual reasoning approaches, which leverage the reasoning abilities of large language models (LLMs) with integrated tools to solve problems, have shown promise as more effective strategies than end-to-end VR methods. However, these approaches face limitations, as frozen LLMs lack tool awareness in VR, leading to performance bottlenecks. While leveraging LLMs for reasoning is widely used in other domains, they are not directly applicable to VR due to limited training data, imperfect tools that introduce errors and reduce data collection efficiency in VR, and challenging in fine-tuning on noisy workflows. To address these challenges, we propose DWIM: i) Discrepancy-aware training Workflow generation, which assesses tool usage and extracts more viable workflows for training; and ii) Instruct-Masking fine-tuning, which guides the model to only clone effective actions, enabling the generation of more practical solutions. Our experiments demonstrate that DWIM achieves state-of-the-art performance across various VR tasks, exhibiting strong generalization on multiple widely-used datasets.

视觉推理工具调用大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。