arXiv:2606.12195cs.CV2026-06被引 2

让视频模型像人一样思考:用闭环推理处理长时视频任务

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

论文配图:InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
图 1 · 摘自论文原文
  • 构建多模态上下文闭环,融合观察、指令与记忆持续推理
  • 在Video-MME等数据集上超越现有模型,支持长视频证据积累与验证
  • 适合做视频理解、智能代理的开发者,尤其关注视觉决策的场景

近期基础模型研究转向多步骤推理与工具使用的智能体行为,但开源工作多聚焦文本主导场景,长期多模态任务仍待探索。视频任务需持续时间理解与迭代交互,该空白尤为明显。本文提出InternVideo3,通过多模态上下文推理(MCR)增强此类能力。MCR将理解视为共享动态上下文中的闭合回路过程,包含观测、指令、推理、工具操作与记忆。这使长视频理解转化为证据累积与验证过程。为保障效率,引入多模态多头隐状态注意力(M^2LA),在保持完整词元流的同时压缩键值缓存。训练采用分阶段策略:持续预训练、短到长监督微调、基于规则的强化学习及在线策略蒸馏。实验显示,InternVideo3在Video-MME、MLVU和EgoSchema等基准上表现优异。进一步将其实例化为带检索工具的视频智能体,展现稳健的证据驱动行为。结果表明,高效上下文处理与闭合回路推理是实现开放多模态模型向长时视觉智能体演进的关键。

原文摘要 · Abstract (English)

Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant settings, leaving long-horizon multimodal tasks underexplored. This gap is evident in video tasks requiring sustained temporal understanding and iterative interaction. We present InternVideo3, a framework enhancing these capabilities via Multimodal Contextual Reasoning (MCR). MCR treats understanding as a closed-loop process over a shared, evolving context containing observations, instructions, reasoning, tool actions, and memory. This frames long-video understanding as evidence accumulation and verification. To ensure efficiency, we introduce Multimodal Multi-head Latent Attention (M^2LA), a token-preserving reparameterization compressing KV-cache states while retaining the full token stream. Our staged training includes continued pretraining, short-to-long supervised fine-tuning, rule-based reinforcement learning, and on-policy distillation. Experiments show InternVideo3 achieves strong performance on benchmarks like Video-MME, MLVU, and EgoSchema. We further instantiate the model as a video agent with retrieval tools, demonstrating robust evidence-grounded behavior. Our results suggest that efficient context handling and closed-loop reasoning are vital for adapting open multimodal models toward long-horizon visually grounded agency.

多模态视频理解智能体推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。