VideoChat3以40亿参数实现高效通用视频理解,全面开源。
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

- 采用膨胀3D视觉变压器与自适应帧分辨率提升效率
- 构建三大高质量数据集,覆盖通用/长视频/流媒体场景
- 模型仅40亿参数却超越更大规模开源模型
近期视频理解研究涵盖运动建模、长视频处理与流式交互,推动该领域向实际应用发展。然而,当前开源模型仍存在诸多局限:泛化能力弱,仅适用于特定领域;计算开销高,影响效率与扩展性;多数模型仅部分开源,训练代码、策略或数据集缺失,阻碍复现与社区协作。为此,我们提出VideoChat3,一个完全开源、高效且通用的视频中心多模态大模型。通过两项互补设计提升性能:为提升效率,引入膨胀3D视觉变换器(I3D-ViT)与自适应帧分辨率机制,实现高效的时空表征并降低训练推理成本;为增强效果,构建可扩展的数据合成流程,生成三个多样化高质量数据集:VideoChat3-Academic2M、VideoChat3-LV116K与VideoChat3-OL617K,分别覆盖通用、长视频与流媒体场景,显著提升模型跨域泛化能力。整合上述设计后,VideoChat3在通用、长视频与流式基准测试中,以仅40亿参数达到优于或等同于更大参数量开源模型的性能,实现了广泛泛化与高效率的罕见平衡。
原文摘要 · Abstract (English)
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。