arXiv:2412.12791cs.CVcs.AI2024-12AAAI被引 10

通过互补掩码实现视频事件定位与描述的隐式对齐,简化流程且效果更优。

Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning

  • 用互补掩码生成正负区域,隐式对齐事件位置与描述
  • 在没有边界标注下,性能超越现有弱监督方法
  • 适合做视频理解、弱监督任务的研究者参考

弱监督密集视频描述(WSDVC)旨在无需事件边界标注的情况下定位并描述视频中所有重要事件。由于缺乏明确的位置监督,准确定位事件时间位置极具挑战。现有方法依赖显式的事件位置与描述之间的对齐约束,训练和推理中需复杂的事例提议过程。为此,我们提出一种基于互补掩码的隐式位置-描述对齐新范式,简化了事件提议与定位流程,同时保持有效性。模型由双模视频描述模块和掩码生成模块组成:前者捕捉全局事件信息并生成描述,后者生成可微分的正负掩码以定位事件。通过确保正负掩码视频生成的描述互补,形成完整描述,从而在弱监督下实现事件位置与描述的隐式对齐。大量实验表明,该方法优于现有弱监督方法,并达到与全监督方法相当的性能。

原文摘要 · Abstract (English)

Weakly-Supervised Dense Video Captioning (WSDVC) aims to localize and describe all events of interest in a video without requiring annotations of event boundaries. This setting poses a great challenge in accurately locating the temporal location of event, as the relevant supervision is unavailable. Existing methods rely on explicit alignment constraints between event locations and captions, which involve complex event proposal procedures during both training and inference. To tackle this problem, we propose a novel implicit location-caption alignment paradigm by complementary masking, which simplifies the complex event proposal and localization process while maintaining effectiveness. Specifically, our model comprises two components: a dual-mode video captioning module and a mask generation module. The dual-mode video captioning module captures global event information and generates descriptive captions, while the mask generation module generates differentiable positive and negative masks for localizing the events. These masks enable the implicit alignment of event locations and captions by ensuring that captions generated from positively and negatively masked videos are complementary, thereby forming a complete video description. In this way, even under weak supervision, the event location and event caption can be aligned implicitly. Extensive experiments on the public datasets demonstrate that our method outperforms existing weakly-supervised methods and achieves competitive results compared to fully-supervised methods.

视频描述弱监督掩码机制事件定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。