arXiv:2508.03284cs.AI2025-08ICCV被引 17

构建真实场景下多步推理的工具型视觉问答数据集

ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools

  • 用深度优先搜索生成人类式工具使用逻辑的多步推理任务
  • 含2.78步平均推理长度,覆盖10种工具7大领域共2.3万实例
  • 小模型在真实场景测试中超越GPT-3.5-turbo,适合通用工具应用研究

将外部工具融入大型基础模型(LFMs)已成为提升其问题求解能力的有前景方法。尽管现有研究在工具增强的视觉问答(VQA)中表现良好,但最新基准显示,在需要多步推理的真实世界复杂多模态场景中仍存在显著性能差距。为此,我们提出ToolVQA,一个包含2.3万实例的大规模多模态数据集,旨在填补这一空白。与依赖合成场景和简化问题的先前数据集不同,ToolVQA采用真实视觉上下文和具有挑战性的隐式多步推理任务,更贴近实际用户交互。我们提出ToolEngine数据生成流水线,结合深度优先搜索(DFS)与动态上下文示例匹配机制,模拟人类式工具使用推理过程。该数据集涵盖10种多模态工具,覆盖7个不同任务领域,平均每例推理步骤为2.78步。在ToolVQA上微调的7B LFM不仅在测试集上表现优异,还在多种分布外(OOD)数据集上超越闭源大模型GPT-3.5-turbo,展现出强大的真实世界工具使用泛化能力。

原文摘要 · Abstract (English)

Integrating external tools into Large Foundation Models (LFMs) has emerged as a promising approach to enhance their problem-solving capabilities. While existing studies have demonstrated strong performance in tool-augmented Visual Question Answering (VQA), recent benchmarks reveal significant gaps in real-world tool-use proficiency, particularly in functionally diverse multimodal settings requiring multi-step reasoning. In this work, we introduce ToolVQA, a large-scale multimodal dataset comprising 23K instances, designed to bridge this gap. Unlike previous datasets that rely on synthetic scenarios and simplified queries, ToolVQA features real-world visual contexts and challenging implicit multi-step reasoning tasks, better aligning with real user interactions. To construct this dataset, we propose ToolEngine, a novel data generation pipeline that employs Depth-First Search (DFS) with a dynamic in-context example matching mechanism to simulate human-like tool-use reasoning. ToolVQA encompasses 10 multimodal tools across 7 diverse task domains, with an average inference length of 2.78 reasoning steps per instance. The fine-tuned 7B LFMs on ToolVQA not only achieve impressive performance on our test set but also surpass the large close-sourced model GPT-3.5-turbo on various out-of-distribution (OOD) datasets, demonstrating strong generalizability to real-world tool-use scenarios.

视觉问答多步推理工具增强数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。