arXiv:2609.04183cs.CVcs.AI2026-09

用视觉语言模型动态生成过渡描述,提升弱监督视频字幕定位精度

Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

论文配图:Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning
图 1 · 摘自论文原文
  • 基于视觉语言模型生成帧级叙事,动态识别事件转换点
  • 通过语义变化点优化时间掩码,使描述与视觉内容对齐
  • 在ActivityNet和YouCook2上达到当前最佳性能,适合视频理解研究者

弱监督密集视频字幕任务旨在仅凭每段视频的事件级标题序列,定位并描述未剪辑视频中的多个事件。现有方法通过大语言模型合成辅助过渡描述以增强视觉-语言对齐,但这些描述缺乏视觉依据,且固定分配于每个事件间隙,位置和时长不变。为此,我们提出Seeing Before Synthesizing(SBS)框架,仅在必要时自适应地提供视觉引导的语义描述。利用视觉语言模型生成事件间隙的帧级叙述,并从其语义变化中检测转换点。针对识别出的转换点,通过融合时间中点与语义变化点,并选择最大化视觉-语言对齐的宽度来精炼事件间时间掩码。在ActivityNet Captions和YouCook2数据集上的实验表明,该方法在字幕生成和定位任务上均达到当前最优性能。

原文摘要 · Abstract (English)

Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted. Leveraging a VLM, we generate frame-level narratives for the inter-event gaps and detect transitions from the semantic variation across them. For identified transitions, we then refine inter-event temporal masks by blending the temporal midpoint with the semantic change point and selecting the width that maximizes vision-language alignment. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance in both captioning and localization.

视频字幕视觉语言模型弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。