让搜索模型真正看图说话,支持全程图文交互与错误恢复。
WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

- 引入原生图文交互框架,图像持久存储可全程引用。
- 在150个任务上提升平均得分19.22点,超越同规模开源模型。
- 适合需要精准视觉理解的长程多模态任务研究者使用。
多模态搜索代理通过整合开放网络中新出现的长尾证据,扩展了参数化知识。然而,现有代理环境常仅以文本形式呈现检索结果,忽略工具返回的图像,导致视觉驱动的推理路径退化为纯文本推理。长时交互还会放大工具调用、响应长度、超时和预算失败问题,可能丢弃仍可挽救的轨迹,浪费计算资源,并干扰策略更新。为此,我们提出WeAgent-Harness,一个支持原生文本-视觉交互与运行时恢复的多模态代理框架。检索图像被赋予持久磁盘引用,使模型可在整个推理轨迹中反复查看、处理并引用。基于此框架,我们构建了WeAgent-MMSearch系统,覆盖数据构建、代理后训练及多模态推演全过程。在数据构建阶段,强大型多模态大模型利用WeAgent-Harness发现、合成并验证多模态搜索任务,收集专家轨迹。后训练阶段,我们提出故障感知的广义策略优化(FA-GSPO),可恢复可挽救的异常轨迹并过滤无效轨迹,提升受限多模态规划与搜索能力。我们还引入VisTarget-Bench,一个包含150个经人工验证的任务基准,每个问题配有一个保留目标图像,用于区分图像检索失败与视觉感知失败。在VisTarget-Bench及七个公开基准上的评估表明,代理后训练使平均得分提升19.22分,使我们的模型超越同等规模的开源模型,并媲美参数量约为其十倍的竞品模型。
原文摘要 · Abstract (English)
Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned images from subsequent context, reducing visually grounded trajectories to text-only reasoning. Long-horizon interaction also compounds tool-call, response-length, timeout, and budget failures, which can discard salvageable trajectories, waste rollout computation, and disturb policy updates. To address these issues, we introduce WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery. Retrieved images receive persistent disk references, allowing the model to inspect, process, and cite them throughout the trajectory. Based on this harness, we develop WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout. For data construction, a strong MLLM uses WeAgent-Harness to discover, synthesize, and verify MMSearch-style tasks and collect expert trajectories. During post-training, our Failure-Aware GSPO (FA-GSPO) recovers salvageable abnormal rollouts and filters invalid ones to improve bounded multimodal planning and search. We also introduce VisTarget-Bench, a 150-task human-verified benchmark that pairs each question with a held-out target image, distinguishing image-retrieval failures from visual-perception failures. Evaluation on VisTarget-Bench and seven public benchmarks shows that agentic post-training improves the average score by 19.22 points, enabling our model to outperform similarly sized open-source models and rival models with roughly ten times its parameter count.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。