让AI在图文混合场景中精准搜索并推理,提升信息获取能力。
M$^3$Searcher: Modular Multimodal Information Seeking Agency with Retrieval-Oriented Reasoning
- 分模块设计,先检索后推理,避免多模态任务混淆
- 在复杂图文任务中准确率超现有方法,支持跨任务迁移
- 专为多跳图文搜索设计新数据集,助力强化学习训练
近年来,基于DeepResearch的智能体在自主获取和整合真实网络环境中的信息方面展现出强大能力。然而,现有方法仍局限于文本模态。将自主信息搜索智能体扩展至多模态场景面临两大挑战:大规模多模态工具使用训练中的专用性与泛化性权衡问题,以及复杂、多步骤多模态搜索轨迹训练数据严重稀缺。为此,我们提出M$^3$Searcher,一种模块化的多模态信息搜寻智能体,通过显式解耦信息获取与答案推导过程实现优化。该智能体采用面向检索的多目标奖励机制,联合优化事实准确性、推理合理性与检索保真度。同时,我们构建了MMSearchVQA数据集,支持以检索为中心的强化学习训练。实验表明,M$^3$Searcher在复杂多模态任务中显著优于现有方法,具备强泛化适应能力与高效推理性能。
原文摘要 · Abstract (English)
Recent advances in DeepResearch-style agents have demonstrated strong capabilities in autonomous information acquisition and synthesize from real-world web environments. However, existing approaches remain fundamentally limited to text modality. Extending autonomous information-seeking agents to multimodal settings introduces critical challenges: the specialization-generalization trade-off that emerges when training models for multimodal tool-use at scale, and the severe scarcity of training data capturing complex, multi-step multimodal search trajectories. To address these challenges, we propose M$^3$Searcher, a modular multimodal information-seeking agent that explicitly decouples information acquisition from answer derivation. M$^3$Searcher is optimized with a retrieval-oriented multi-objective reward that jointly encourages factual accuracy, reasoning soundness, and retrieval fidelity. In addition, we develop MMSearchVQA, a multimodal multi-hop dataset to support retrieval centric RL training. Experimental results demonstrate that M$^3$Searcher outperforms existing approaches, exhibits strong transfer adaptability and effective reasoning in complex multimodal tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。