构建首个覆盖三类信息流的多模态浏览数据框架,提升智能体泛化能力。
UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp

- 统一生成文本、图像转文本、文本转图像三类浏览数据
- 训练出350亿参数智能体,在5个基准上达54.4%准确率
- 适合研究多模态工具使用与长程推理的开发者和研究人员
多模态浏览任务要求智能体在动态网页内容中融合感知、工具调用与长周期推理,面临组合结构处理、开放世界不确定性及跨模态整合的挑战。真实场景下的多模态浏览包含三种信息流模式:纯文本、图像到文本、文本到图像,但现有数据构建方法仅覆盖前两种,导致文本到图像模式缺失,制约智能体通用性与鲁棒性。我们提出UNIBROWSE,首个统一的数据流水线,首次同时生成涵盖三种模式的训练数据,通过实时网络检索增强知识图谱以提升保真度,并引入探索度新指标过滤低信号样本,实现高效强化学习。基于该流水线,我们生成高质量冷启动工具使用轨迹与探索丰富的问答对,采用监督微调与探索感知强化学习训练350亿参数智能体。结果表明,所提UNIBROWSE智能体在五个多样化基准上达到平均54.4%准确率,较基础模型Qwen3.5-35B-A3B提升10.5个百分点,超越GPT-5(42.9)、Gemini-2.5 Pro(44.8)和Gemini-2.5 Flash(41.3)等闭源工作流。
原文摘要 · Abstract (English)
Multimodal BrowseComp tasks require agents to combine perception, tool use, and long-horizon reasoning over dynamic web content, challenging their ability to handle compositional structure, open-world uncertainty, and multimodal integration across extended interactions. Crucially, real-world multimodal browsing involves three distinct information-flow patterns: text-only, image-to-text, and text-to-image, yet existing data construction methods cover only the text-only and image-to-text patterns, leaving text-to-image largely unaddressed and limiting agent generality and robustness. We introduce UNIBROWSE, a unified data pipeline that for the first time simultaneously generates training data covering all three patterns, augments curated knowledge graphs with live web retrieval for improved fidelity, and introduces a novel metric of exploration degree to filter low-signal instances for efficient reinforcement learning. Through this pipeline, we produce high-quality cold-start tool-use trajectories and exploration-rich QA pairs, and train a 35B-scale agent via supervised fine-tuning and exploration-aware RL.The resulting UNIBROWSE agent achieves state-of-the-art performance on multimodal BrowseComp benchmarks, attaining an average accuracy of 54.4 across five diverse benchmarks -- an improvement of 10.5 points over its base model Qwen3.5-35B-A3B -- and surpassing serveral closed-source agent workflows such as GPT-5 (42.9), Gemini-2.5 Pro (44.8), and Gemini-2.5 Flash (41.3).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。