arXiv:2606.02522cs.CVcs.AI2026-06被引 3

测试视频大模型对瞬间视觉事件的捕捉能力,发现多数模型表现不佳。

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

论文配图:Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
图 1 · 摘自论文原文
  • 设计新基准测试瞬时视觉事件理解能力
  • 顶级模型准确率仅39.6%,多数开源模型低于25%
  • 适合关注视频模型时序精度的研究者和开发者

视频多模态大模型在通用和长视频理解上进展迅速,但对短暂关键视觉证据的保留能力仍缺乏研究。许多实际问题由瞬时视觉事件(如仅持续数帧的动作或状态变化)决定。这类证据可能被稀疏采样遗漏、视觉令牌压缩压制或粗粒度时序聚合稀释,导致错误,语言推理难以挽回。本文提出Moment-Video基准,通过瞬时视觉事件理解诊断视频MLLM的时序保真度。每个问题基于局部、可观察且采样敏感的事件,要求模型识别、计数、描述或推理瞬态证据,而非依赖持久物体、全局场景或语言先验。该数据集包含1,000个经人工验证的视频问答对,覆盖7个领域和25个细粒度子类别,涵盖四种任务类型:时序发生、时序计数、动作描述与时序推理。我们在Moment-Video上评估了33个专有及开源模型。表现最佳的Seed-2.0-Pro模型整体准确率为39.6%,大多数开源模型低于25%,暴露出显著的能力差距。诊断分析表明,更密集的帧采样虽提升部分模型性能,但无法消除瓶颈,更长视频则加剧时序定位挑战。结果表明,当前视频MLLM仍缺乏捕捉、保留和利用短暂关键视觉证据的时序忠实表征。

原文摘要 · Abstract (English)

Video multimodal large language models (MLLMs) have made rapid progress on general and long-form video understanding, yet their ability to preserve brief answer-critical visual evidence remains underexplored. Many practical questions are determined by momentary visual events: localized actions or state transitions that may last only a few frames. Such evidence can be skipped by sparse frame sampling, suppressed by visual-token compression, or diluted by coarse temporal aggregation, causing failures that language-side reasoning cannot reliably recover. We introduce Moment-Video, a benchmark for diagnosing the temporal fidelity of video MLLMs through momentary visual event understanding. Each question is grounded in a localized, visually observable, and sampling-sensitive event, requiring models to notice, count, describe, or reason about transient evidence rather than rely on persistent objects, global scene context, or language priors. Moment-Video contains 1,000 human-verified video-QA pairs across 7 domains and 25 fine-grained subcategories, covering four task types: Temporal Occurrence, Temporal Counting, Action Description, and Temporal Reasoning. We evaluate 33 proprietary and open-source MLLMs on Moment-Video. The best-performing model, Seed-2.0-Pro, achieves only 39.6% overall accuracy, while most open-source models remain below 25%, revealing a substantial gap in momentary visual event understanding. Diagnostic analyses show that denser frame sampling improves some models but does not eliminate the bottleneck, and longer videos introduce stronger temporal-localization challenges. These findings suggest that current video MLLMs still lack temporally faithful representations for capturing, preserving, and using brief but decisive visual evidence.

视频理解时序精度多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。