arXiv:2512.02395cs.CV2025-12被引 26

通过视觉与搜索交替思考,实现无需强化学习的多模态智能体推理

Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch

论文配图:Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
图 1 · 摘自论文原文
  • 采用视觉操作与外部知识检索交替进行的混合推理策略
  • 在MMSearch和FVQA上分别达到66.1和67.2分,超越Gemini 2.5 Flash所有11项指标
  • 仅用监督微调即可实现超过10步复杂任务的长程规划,适合多模态智能体研究者

尽管多模态智能体系统取得进展,现有方法常将图像处理与网络搜索视为独立能力,严重依赖昂贵的强化学习,并缺乏基于真实工具执行轨迹的规划。为此,我们提出Skywork-R1V4,一个300亿参数(A3B)的多模态智能体模型,统一了多模态规划、主动图像操作(“以图思考”)、深度多模态搜索及关键的交替推理机制——动态在视觉操作与外部知识获取间切换。该模型仅在少于3万条高质量、规划-执行一致的轨迹上通过监督微调训练,并经逐步一致性过滤验证。其在感知与多模态搜索基准上达到领先表现:MMSearch得分为66.1,FVQA得分为67.2,全面超越Gemini 2.5 Flash在全部11个评测指标上的结果。模型在推理时展现出涌现的长程推理能力,成功协调超过10次工具调用以解决复杂多步任务。结果表明,通过精心构建的监督学习即可实现高级多模态智能体能力,无需依赖强化学习。

原文摘要 · Abstract (English)

Despite recent progress in multimodal agentic systems, existing approaches often treat image manipulation and web search as disjoint capabilities, rely heavily on costly reinforcement learning, and lack planning grounded in real tool-execution traces. To address these limitations, we present Skywork-R1V4, a 30B (A3B) parameter multimodal agentic model that unifies multimodal planning, active image manipulation ("thinking with images"), deep multimodal search, and, most critically, interleaved reasoning that dynamically alternates between visual operations and external knowledge retrieval. Trained solely via supervised fine-tuning on fewer than 30,000 high-quality, planning-execution-consistent trajectories and validated through stepwise consistency filtering, Skywork-R1V4 achieves state-of-the-art results across perception and multimodal search benchmarks: it scores 66.1 on MMSearch and 67.2 on FVQA, surpassing Gemini 2.5 Flash on all 11 metrics. Skywork-R1V4 exhibits emergent long-horizon reasoning at inference time, successfully orchestrating more than 10 tool calls to solve complex, multi-step tasks. Our results demonstrate that sophisticated agentic multimodal intelligence can be achieved through carefully curated supervised learning alone, without any reliance on reinforcement learning.

多模态智能体交替推理监督学习图像思考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。