LLaVA-OneVision-2通过动态编码流实现高效长视频理解,性能全面超越现有模型。
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

- 将压缩视频转为比特成本流,动态分组并聚焦关键时空信息
- 在JumpScore上达74.9分,比Qwen3-VL高44.8点,时序定位提升9.7点
- 适合需要精准视频理解与多模态推理的研究者与开发者
我们提出LLaVA-OneVision-2(LLaVA-OV-2),该系列迄今最强大的视觉语言模型,在多种多模态基准上表现卓越。模型基于原生OneVision-Encoder,采用窗口化注意力实现高效局部计算,同时保持原生分辨率。其核心创新是编码器流标记化:将压缩视频视为连续比特成本流,由比特成本动态决定自适应时间分组,运动残差线索选取显著空间特征生成紧凑视觉画布。该分配策略将有限的标记预算集中于承载事件的内容,实现比固定帧组更稳定的长视频标记压缩。共享3D RoPE将编码画布、采样帧与图像统一至时空坐标系。我们还构建了大规模开放监督的数据与训练栈:约800万条重标注视频样本用于预训练,400万条空间语料用于微调。引入JumpScore,一个针对高频密集重复动作中细粒度定位的时序定位基准,填补现有视频评估空白。LLaVA-OV-2展现出统一的视频理解、时序定位、空间定位与操作轨迹推理能力。在JumpScore上,LLaVA-OneVision-2-8B达到74.9跳分mAP,较Qwen3-VL-8B高出44.8点;在相同视觉标记预算下,编码器流输入相比帧采样使时序定位提升9.7点。在标准基准上,该模型在视频任务上领先Qwen3-VL-8B 4.3分,空间任务领先5.3分,跟踪任务平均J&F提升15.6分。
原文摘要 · Abstract (English)
We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of multimodal benchmarks. The model builds on a native OneVision-Encoder and incorporates Windowed Attention for efficient local computation while maintaining native resolution. Its key advance is codec-stream tokenization: it treats compressed video as a continuous bit-cost stream, where bit-cost dynamics determine adaptive temporal groups, and motion-residual cues select salient spatial evidence into compact visual canvases. This allocation concentrates a limited token budget on event-bearing content, enabling more stable long-video token compression than fixed groups of pictures. A shared 3D RoPE further places codec canvases, sampled frames, and images in a unified spatiotemporal coordinate system. Furthermore, we build the LLaVA-OV-2 data and training stack around large-scale open supervision: approximately 8M re-captioned video samples for pretraining, a 4M-sample spatial corpus for fine-tuning. We also introduce JumpScore, a temporal-localization benchmark targeting fine-grained grounding in high-frequency, densely repeated motion, a regime underrepresented by existing video evaluations. A standout capability of LLaVA-OV-2 is its unified perception across video understanding, temporal grounding, spatial grounding, and manipulation-trace reasoning. On JumpScore, LLaVA-OneVision-2-8B reaches 74.9 JumpScore mAP, surpassing Qwen3-VL-8B (30.1) by +44.8 points; under matched visual-token budgets on the same benchmark, codec-stream inputs improve temporal grounding over frame sampling by +9.7 points. Across standard benchmarks, LLaVA-OneVision-2-8B further outperforms Qwen3-VL-8B by +4.3 average points on video tasks, +5.3 on spatial tasks, and +15.6 average J&F on tracking tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。