arXiv:2508.05748cs.IR2025-08被引 112

WebWatcher让AI能看图搜信息,解决复杂跨模态问题

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

  • 用合成多模态数据训练,快速启动视觉语言推理能力
  • 在4个视觉问答基准上超越主流模型和开源代理
  • 适合研究多模态智能体、信息检索的开发者与学者

如Deep Research等网络智能体已展现超人类认知能力,能解决高难度信息查询任务。但现有研究多聚焦文本,忽视真实世界中的视觉信息,导致多模态深度研究极具挑战——这类智能体需更强的感知、逻辑、知识运用与工具调用能力。为此,我们提出WebWatcher,一个具备增强视觉-语言推理能力的多模态深度研究智能体。它利用高质量合成多模态轨迹实现高效冷启动训练,结合多种工具进行深度推理,并通过强化学习提升泛化性能。为更准确评估多模态智能体能力,我们构建了BrowseComp-VL基准,其设计类似BrowseComp,要求同时处理视觉与文本信息的复杂信息检索任务。实验表明,WebWatcher在四个挑战性VQA基准上显著优于专有基线、RAG工作流及开源智能体,为解决复杂多模态信息查询任务开辟新路径。

原文摘要 · Abstract (English)

Web agents such as Deep Research have demonstrated superhuman cognitive abilities, capable of solving highly challenging information-seeking problems. However, most research remains primarily text-centric, overlooking visual information in the real world. This makes multimodal Deep Research highly challenging, as such agents require much stronger reasoning abilities in perception, logic, knowledge, and the use of more sophisticated tools compared to text-based agents. To address this limitation, we introduce WebWatcher, a multi-modal Agent for Deep Research equipped with enhanced visual-language reasoning capabilities. It leverages high-quality synthetic multimodal trajectories for efficient cold start training, utilizes various tools for deep reasoning, and further enhances generalization through reinforcement learning. To better evaluate the capabilities of multimodal agents, we propose BrowseComp-VL, a benchmark with BrowseComp-style that requires complex information retrieval involving both visual and textual information. Experimental results show that WebWatcher significantly outperforms proprietary baseline, RAG workflow and open-source agents in four challenging VQA benchmarks, which paves the way for solving complex multimodal information-seeking tasks.

多模态智能体视觉问答信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。