用自强化机制让视频模型仅凭问答对学会视觉隐空间推理,效率提升68倍。
VideoLatent: Video-Language Learning via Latent Self-Forcing

- 通过隐空间自强化训练,仅用标准三元组数据学习视频隐表示。
- 在14个基准上超越现有模型,训练和推理开销分别降低6倍和68倍。
- 适配不同模型结构与规模,适合追求高效视频理解的研究者。
链式思维(CoT)推理虽能提升多模态大语言模型(MLLM)的视频理解能力,但需大量人工标注且计算开销大。现有视觉隐空间推理方法多针对图像任务,依赖额外监督信号(如CoT轨迹、辅助图像或细粒度标注),难以拓展至视频场景。为此,我们提出VideoLatent,一种专为视频理解设计的MLLM,其核心是隐空间注入模块与新颖的隐空间自强化训练范式,包含隐空间对齐与多样性目标,仅需标准视频-问题-答案三元组即可训练。在14个基准上的实验证明,该模型在通用视频理解与复杂推理任务中持续优于现有标准及隐空间模型。相比Video-R1,训练/推理开销分别减少约6倍/约68倍。此外,实验表明该方法对不同MLLM主干网络和模型规模具有强泛化能力。
原文摘要 · Abstract (English)
Recent advancements in chain-of-thought (CoT) reasoning have shown promise in enhancing video understanding and reasoning capabilities of multimodal large language models (MLLMs). However, existing CoT-based MLLMs require labor-intensive CoT annotations and incur substantial training and inference overhead. While visual latent reasoning has emerged as a more efficient alternative, existing methods primarily focus on image tasks and heavily rely on additional supervision signals for visual latent generation (e.g., CoT traces, auxiliary images, or fine-grained annotations), limiting their scalability and transferability to video tasks. To bridge this gap, we introduce VideoLatent, a novel MLLM equipped with a latent injection module tailored for video understanding and reasoning. Specifically, VideoLatent learns to perform visual latent reasoning using a new latent self-forcing training paradigm, which comprises latent alignment and latent diversity objectives, and relies solely on standard video-question-answer triplets. Extensive experiments across 14 benchmarks demonstrate that our model consistently outperforms existing standard and latent MLLMs on general video understanding and complex video reasoning. Compared with Video-R1, our VideoLatent achieves superior computational efficiency, reducing training/inference overhead by $\sim$6$\times$/$\sim$68$\times$. Moreover, experiments demonstrate that our method has strong generalizability to different MLLM backbones and different model scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。