arXiv:2601.02908cs.CVcs.AI2026-01中稿 · WACV 2026被引 1

用时间锚点提升视频大模型对事件边界的精准定位能力

TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors

  • 引入时间锚点学习事件定位,引导大模型进行时序感知理解
  • 在多个基准数据集上超越现有方法,显著提升事件检索与时序问答性能
  • 适合关注视频理解、时序定位与多模态生成的研究者

密集视频字幕任务旨在对输入视频中所有时序定位的事件进行解释与描述。近期先进方法利用大语言模型(LLMs)为视频数据提供详细事件描述,但现有视频大模型在未剪辑视频中仍难以精确定位事件边界,导致生成字幕缺乏准确锚定。本文提出TA-Prompting,通过时间锚点学习精确事件定位,并引导视频大模型实现时序感知的事件理解。推理时,为从任意数量事件中生成连贯字幕序列,我们设计了一种事件一致采样策略,选择在时间事件间具有足够连贯性且与视频跨模态相似度高的字幕。在多个基准数据集上的大量实验表明,TA-Prompting优于当前最先进视频大模型,在密集视频字幕、事件定位与时序问答任务中均取得更优表现。

原文摘要 · Abstract (English)

Dense video captioning aims to interpret and describe all temporally localized events throughout an input video. Recent state-of-the-art methods leverage large language models (LLMs) to provide detailed moment descriptions for video data. However, existing VideoLLMs remain challenging in identifying precise event boundaries in untrimmed videos, causing the generated captions to be not properly grounded. In this paper, we propose TA-Prompting, which enhances VideoLLMs via Temporal Anchors that learn to precisely localize events and prompt the VideoLLMs to perform temporal-aware video event understanding. During inference, in order to properly determine the output caption sequence from an arbitrary number of events presented within a video, we introduce an event coherent sampling strategy to select event captions with sufficient coherence across temporal events and cross-modal similarity with the given video. Through extensive experiments on benchmark datasets, we show that our TA-Prompting is favorable against state-of-the-art VideoLLMs, yielding superior performance on dense video captioning and temporal understanding tasks including moment retrieval and temporalQA.

视频理解时序定位大模型字幕生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。