arXiv:2509.09263cs.CV2025-09被引 10

让大模型精准理解长视频时间线,提升事件定位能力。

DATE: Dynamic Absolute Time Enhancement for Long Video Understanding

  • 用时间戳嵌入构建连续时间参考系统
  • 通过语义相关采样确保关键事件覆盖
  • 7B模型性能超越多数72B模型

长视频理解仍是多模态大模型的核心挑战,尤其在需要精确时序推理和事件定位的任务中。现有方法通常采用均匀帧采样并依赖隐式位置编码来建模时间顺序,但难以处理长距离依赖,导致关键信息丢失和时序理解能力下降。本文提出动态绝对时间增强(DATE),通过时间戳注入机制(TIM)与语义引导的时序感知相似性采样(TASS)策略,增强多模态大模型的时序感知能力。具体地,将视频帧嵌入与文本时间戳标记交错,构建连续时间参考体系;并将视频采样重构为视觉-语言检索任务,采用两阶段算法:先将查询扩展为描述性标题以更好对齐视觉特征,再使用相似度驱动的时序正则化贪心策略采样关键事件。该方法在小时级视频基准测试中显著提升绝对时间理解和关键事件定位性能,在7B与72B模型中均达到领先水平,尤其7B模型在部分任务上超越多数72B模型。

原文摘要 · Abstract (English)

Long video understanding remains a fundamental challenge for multimodal large language models (MLLMs), particularly in tasks requiring precise temporal reasoning and event localization. Existing approaches typically adopt uniform frame sampling and rely on implicit position encodings to model temporal order. However, these methods struggle with long-range dependencies, leading to critical information loss and degraded temporal comprehension. In this paper, we propose Dynamic Absolute Time Enhancement (DATE) that enhances temporal awareness in MLLMs through the Timestamp Injection Mechanism (TIM) and a semantically guided Temporal-Aware Similarity Sampling (TASS) strategy. Specifically, we interleave video frame embeddings with textual timestamp tokens to construct a continuous temporal reference system. We further reformulate the video sampling problem as a vision-language retrieval task and introduce a two-stage algorithm to ensure both semantic relevance and temporal coverage: enriching each query into a descriptive caption to better align with the vision feature, and sampling key event with a similarity-driven temporally regularized greedy strategy. Our method achieves remarkable improvements w.r.t. absolute time understanding and key event localization, resulting in state-of-the-art performance among 7B and 72B models on hour-long video benchmarks. Particularly, our 7B model even exceeds many 72B models on some benchmarks.

视频理解时序建模大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。