arXiv:2603.05437cs.CVcs.AI2026-03中稿 · CVPR

通过语义对齐与生成增广,提升弱监督视频描述的定位精度。

SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

  • 基于跨模态对齐构建语义感知掩码,强化视频区域与描述的关联。
  • 在ActivityNet和YouCook2上达到最优的生成与定位性能。
  • 利用大模型生成合成描述,缓解标注稀疏问题,适合弱监督研究者。

弱监督密集视频描述旨在仅使用描述注释而无需时间边界信息的情况下定位并描述视频中的事件。现有方法基于高斯掩码和互补描述引入隐式监督,但仅关注生成非重叠掩码,未考虑其与对应事件的语义关系,导致掩码分布简单、缺乏语义意义。此外,依赖真实描述造成性能受限,因数据集本身标注稀疏。本文提出SAIL,通过跨模态对齐构建语义感知掩码,相似性感知训练目标使掩码聚焦于与事件描述高度相似的视频区域。为在标注稀疏场景下实现更精确的掩码生成,引入基于大语言模型的增广策略,生成合成描述作为额外对齐信号。这些合成描述通过跨掩码机制融入训练,提供辅助引导以提升定位精度,且不损害主目标。在ActivityNet Captions和YouCook2数据集上的实验表明,该方法在生成与定位指标上均达到当前最优性能。

原文摘要 · Abstract (English)

Weakly-Supervised Dense Video Captioning aims to localize and describe events in videos trained only on caption annotations, without temporal boundaries. Prior work introduced an implicit supervision paradigm based on Gaussian masking and complementary captioning. However, existing method focuses merely on generating non-overlapping masks without considering their semantic relationship to corresponding events, resulting in simplistic, uniformly distributed masks that fail to capture semantically meaningful regions. Moreover, relying solely on ground-truth captions leads to sub-optimal performance due to the inherent sparsity of existing datasets. In this work, we propose SAIL, which constructs semantically-aware masks through cross-modal alignment. Our similarity aware training objective guides masks to emphasize video regions with high similarity to their corresponding event captions. Furthermore, to guide more accurate mask generation under sparse annotation settings, we introduce an LLM-based augmentation strategy that generates synthetic captions to provide additional alignment signals. These synthetic captions are incorporated through an inter-mask mechanism, providing auxiliary guidance for precise temporal localization without degrading the main objective. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance on both captioning and localization metrics.

视频描述弱监督语义对齐数据增广

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。