arXiv:2509.04602cs.CV2025-09EMNLP被引 7

通过关注视频关键帧和场景变化,提升密集视频字幕生成效果

Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning

  • 根据时间戳生成帧重要性权重,让模型聚焦关键帧
  • 按帧相似性分割视频,更好捕捉场景切换点
  • 在YouCook2和ViTT数据集上达到当前最优性能

密集视频字幕任务旨在定位视频中的事件并为每个事件生成描述。现有端到端方法存在两个问题:(1) 仅对文本施加时间戳监督,忽略视频帧差异;(2) 从固定长度视频块中检索字幕,难以捕捉场景转换。为此,我们提出Sali4Vid,一种简单有效的显著性感知框架。引入显著性感知视频重加权,将时间戳标注转化为基于sigmoid的帧重要性权重;提出语义自适应字幕检索,通过帧相似性分割视频以捕捉场景转换,提升字幕召回能力。Sali4Vid在YouCook2和ViTT数据集上取得当前最优结果,证明联合优化视频加权与检索对密集视频字幕的有效性。

原文摘要 · Abstract (English)

Dense video captioning aims to temporally localize events in video and generate captions for each event. While recent works propose end-to-end models, they suffer from two limitations: (1) applying timestamp supervision only to text while treating all video frames equally, and (2) retrieving captions from fixed-size video chunks, overlooking scene transitions. To address these, we propose Sali4Vid, a simple yet effective saliency-aware framework. We introduce Saliency-aware Video Reweighting, which converts timestamp annotations into sigmoid-based frame importance weights, and Semantic-based Adaptive Caption Retrieval, which segments videos by frame similarity to capture scene transitions and improve caption retrieval. Sali4Vid achieves state-of-the-art results on YouCook2 and ViTT, demonstrating the benefit of jointly improving video weighting and retrieval for dense video captioning

视频字幕显著性感知场景分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。