arXiv:2508.15641cs.CV2025-08被引 6

让视频大模型精准定位事件时间与实体互动,提升长视频理解能力

When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding

  • 用扩散模型思想建模时间特征,增强事件边界感知
  • 通过实体定位显式绑定视觉证据,实现跨时精准对齐
  • 引入离散时间标记,支持细粒度时间推理,适合复杂视频任务

理解视频不仅需回答开放问题,更需精确定位事件发生时间及实体间跨时交互。现有视频大模型虽在整体推理上进展显著,但在时间感知上仍较粗略:时间戳仅隐式编码,帧级特征难以捕捉连续性,语言与视觉对齐常偏离关注实体。本文提出Grounded VideoDiT,通过三项创新克服上述局限:第一,扩散时间潜空间(DTL)编码器增强边界敏感性并保持时间一致性;第二,物体锚定表示将查询实体显式关联到局部视觉证据,强化对齐;第三,混合标记方案结合离散时间标记,实现显式时间建模,支持细粒度时间推理。实验在Charades STA、NExT GQA及多个VideoQA基准上均达到顶尖水平,验证了其强大的定位能力。

原文摘要 · Abstract (English)

Understanding videos requires more than answering open ended questions, it demands the ability to pinpoint when events occur and how entities interact across time. While recent Video LLMs have achieved remarkable progress in holistic reasoning, they remain coarse in temporal perception: timestamps are encoded only implicitly, frame level features are weak in capturing continuity, and language vision alignment often drifts from the entities of interest. In this paper, we present Grounded VideoDiT, a Video LLM designed to overcome these limitations by introducing three key innovations. First, a Diffusion Temporal Latent (DTL) encoder enhances boundary sensitivity and maintains temporal consistency. Second, object grounded representations explicitly bind query entities to localized visual evidence, strengthening alignment. Third, a mixed token scheme with discrete temporal tokens provides explicit timestamp modeling, enabling fine grained temporal reasoning. Together, these designs equip Grounded VideoDiT with robust grounding capabilities, as validated by state of the art results on Charades STA, NExT GQA, and multiple VideoQA benchmarks.

视频理解时间定位实体对齐扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。