arXiv:2601.10611cs.CVcs.AI2026-01被引 122

开源视频语言模型Molmo2实现高精度像素级定位,填补开放数据空白。

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

  • 自建7个视频与2个多图数据集,不依赖闭源模型生成数据
  • 8B模型在短视频任务中表现最优,视频计数准确率达35.5%
  • 支持图像/视频中的点选定位,超越现有开源及部分闭源模型

当前最强的视频语言模型(VLMs)多为专有模型。现有开源模型或依赖闭源模型生成的合成数据进行蒸馏,或未公开训练数据与方法。这导致开源社区难以突破现有视频与图像语言模型的性能上限。尤其许多下游应用不仅需要高层理解,还需像素级定位(如点选或追踪)。即使闭源模型也普遍缺乏此能力。我们提出Molmo2,一套开源视频语言模型,是目前开源模型中顶尖水平,并在单图、多图及视频任务中展现出卓越的点驱动定位能力。核心贡献包括构建7个新视频数据集和2个多图数据集,涵盖高质量视频描述、自由问答、复杂查询物体追踪及创新视频点选数据,均未使用闭源模型生成。我们还提出高效的数据打包与消息树编码训练方案,采用双向视觉注意力与新型令牌权重策略提升性能。最佳8B模型在短视频任务中优于同类开源模型,在计数与描述任务上领先;在长视频任务中表现竞争力。在视频定位任务中,显著超越Qwen3-VL(35.5 vs 29.6准确率),并在视频点选(38.4 vs 20.0 F1)和视频追踪(56.2 vs 41.1 J&F)上超越闭源模型Gemini 3 Pro。

原文摘要 · Abstract (English)

Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from them, or do not disclose their training data or recipe. As a result, the open-source community lacks the foundations needed to improve on the state-of-the-art video (and image) language models. Crucially, many downstream applications require more than just high-level video understanding; they require grounding -- either by pointing or by tracking in pixels. Even proprietary models lack this capability. We present Molmo2, a new family of VLMs that are state-of-the-art among open-source models and demonstrate exceptional new capabilities in point-driven grounding in single image, multi-image, and video tasks. Our key contribution is a collection of 7 new video datasets and 2 multi-image datasets, including a dataset of highly detailed video captions for pre-training, a free-form video Q&A dataset for fine-tuning, a new object tracking dataset with complex queries, and an innovative new video pointing dataset, all collected without the use of closed VLMs. We also present a training recipe for this data utilizing an efficient packing and message-tree encoding scheme, and show bi-directional attention on vision tokens and a novel token-weight strategy improves performance. Our best-in-class 8B model outperforms others in the class of open weight and data models on short videos, counting, and captioning, and is competitive on long-videos. On video-grounding Molmo2 significantly outperforms existing open-weight models like Qwen3-VL (35.5 vs 29.6 accuracy on video counting) and surpasses proprietary models like Gemini 3 Pro on some tasks (38.4 vs 20.0 F1 on video pointing and 56.2 vs 41.1 J&F on video tracking).

视频语言模型开源模型视频定位多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。