系统梳理视频大模型后训练方法,助力模型从看懂到推理。
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- 分三类:思维链微调、可验证目标强化学习、测试时扩展推理
- 提升长视频理解与多模态证据融合能力,突破时空定位瓶颈
- 适合研究视频推理的学者,提供完整评估框架和资源
视频理解是计算机视觉最具挑战性的前沿领域,要求模型能推理复杂时空关系、长期依赖及多模态证据。近年来,集成视觉编码器与强大解码器语言模型的视频大模态模型(Video-LMMs)在视频理解任务中展现出卓越能力。然而,将这些模型从基础感知系统转化为高级推理引擎的关键阶段——后训练,仍分散于文献中。本文首次全面审视 Video-LMM 的后训练方法,涵盖三大核心支柱:带思维链的监督微调(SFT)、基于可验证目标的强化学习(RL),以及通过增强推理计算实现的测试时扩展(TTS)。我们构建了结构化分类体系,阐明各类技术的角色、关联及其在视频任务中的特化应用,解决时间定位、时空定位、长视频效率及多模态证据融合等独特挑战。通过对代表性方法的系统分析,提炼关键设计原则、洞见与评估协议,并指出奖励设计、可扩展性与性价比优化等关键开放问题。此外,还整理了重要基准、数据集与指标,以支持后训练效果的严谨评估。本综述旨在为研究人员与实践者提供统一框架,推动 Video-LMM 能力的持续演进。更多资源与更新请访问:https://github.com/yunlong10/Awesome-Video-LMM-Post-Training
原文摘要 · Abstract (English)
Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and multimodal evidence. The recent emergence of Video-Large Multimodal Models (Video-LMMs), which integrate visual encoders with powerful decoder-based language models, has demonstrated remarkable capabilities in video understanding tasks. However, the critical phase that transforms these models from basic perception systems into sophisticated reasoning engines, post-training, remains fragmented across the literature. This survey provides the first comprehensive examination of post-training methodologies for Video-LMMs, encompassing three fundamental pillars: supervised fine-tuning (SFT) with chain-of-thought, reinforcement learning (RL) from verifiable objectives, and test-time scaling (TTS) through enhanced inference computation. We present a structured taxonomy that clarifies the roles, interconnections, and video-specific adaptations of these techniques, addressing unique challenges such as temporal localization, spatiotemporal grounding, long video efficiency, and multimodal evidence integration. Through systematic analysis of representative methods, we synthesize key design principles, insights, and evaluation protocols while identifying critical open challenges in reward design, scalability, and cost-performance optimization. We further curate essential benchmarks, datasets, and metrics to facilitate rigorous assessment of post-training effectiveness. This survey aims to provide researchers and practitioners with a unified framework for advancing Video-LMM capabilities. Additional resources and updates are maintained at: https://github.com/yunlong10/Awesome-Video-LMM-Post-Training
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。