arXiv:2603.02795cs.CV2026-03被引 3

用强化学习让多模态模型变成能长期搜索的智能代理。

VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning

  • 通过强化学习训练多模态模型,实现跨轮次的网页、图像和文本搜索。
  • 构建大规模复杂多模态问答数据集,提升模型在真实环境中的搜索能力。
  • 提出新基准MM-SearchExam,适合评估多模态搜索智能体性能。

大模型正逐渐成为能与现实环境交互并使用外部工具扩展能力的自主智能体。然而,现有研究多集中于纯文本大语言模型,仅支持单一模态,应用场景受限。多模态大模型虽具备更强感知能力,却仍依赖静态知识,无法获取实时网络信息。本文提出VSearcher,通过强化学习将静态多模态模型转化为可在真实网络环境中执行长时程、多轮次工具调用的多模态搜索代理,支持文本搜索、图像搜索和网页浏览。我们设计了迭代注入数据生成流程,构建大规模复杂多模态问答数据集,并通过综合指标筛选确保数据质量与难度。采用SFT-then-RL训练流程,使基础多模态模型具备多轮工具调用能力。此外,提出多模态搜索评测基准MM-SearchExam,对当前主流专有模型构成挑战。多组实验表明,VSearcher在多个多模态搜索基准上表现优异,超越多个现有多模态搜索代理,甚至在部分任务上优于若干专有模型。

原文摘要 · Abstract (English)

Large models are increasingly becoming autonomous agents that interact with real-world environments and use external tools to augment their static capabilities. However, most recent progress has focused on text-only large language models, which are limited to a single modality and therefore have narrower application scenarios. On the other hand, multimodal large models, while offering stronger perceptual capabilities, remain limited to static knowledge and lack the ability to access and leverage up-to-date web information. In this paper, we propose VSearcher, turning static multimodal model into multimodal search agent capable of long-horizon, multi-turn tool use in real-world web environments, including text search, image search, and web browsing, via reinforcement learning. Specifically, we introduce Iterative Injection Data Synthesis pipeline to generate large-scale, complex multimodal QA questions, which are further filtered with comprehensive metrics to ensure high quality and sufficient difficulty. We then adopt an SFT-then-RL training pipeline to turn base multimodal models to agent capable of multi-turn tool calling in real-world web environments. Besides, we propose a multimodal search benchmark MM-SearchExam dedicated to evaluating search capabilities of multimodal search agents, which proves highly challenging for recent proprietary models. Extensive evaluations across multiple multimodal search benchmarks reveal effectiveness of our method. VSearcher achieves superior performance compared to recent multimodal search agents and even surpasses several proprietary models on multimodal web search tasks.

多模态搜索强化学习智能代理网页搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。