让视觉持续参与长程多轮搜索,提升复杂问题的推理能力
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

- 构建多模态事件图生成带视觉依赖的长链问题
- 设计主动获取视觉信息的智能体,支持动态加载图像
- 无需强化学习,直接微调模型即可实现高效视觉驱动搜索
多模态大语言模型虽提升了视觉理解与推理能力,但其静态参数知识难以应对知识密集且动态变化的开放世界问题。为此,多模态深度搜索成为开放世界信息获取的关键方向,正从单轮事实检索演变为以视觉证据引导的长时程、多轮搜索。然而现有方法通常仅将视觉用于输入或答案阶段,忽视其在中间推理中的作用,且缺乏对长时程交互的设计,导致视觉证据难以为后续检索提供持续驱动力,限制了交互深度与推理跨度。为此,我们提出 DeepVoyager-VL,一种面向视觉闭环的长时程多模态深度搜索框架。具体地,我们构建多模态事件图以生成具有中间视觉依赖和长推理链的问题;设计支持主动视觉采集与按需图像加载的智能体框架;并在合成数据上微调模型,无需强化学习。在十个多模态搜索基准上的实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。