通过循环协作提升视频事件检测与描述精度
Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-learning
- 帧级概念检测生成时间线索,增强视频特征
- 生成器与定位器循环互促,实现语义感知与定位双赢
- 在ActivityNet和YouCook2上达领先效果,可解释性强
密集视频描述旨在检测并描述未剪辑视频中的所有事件。本文提出多概念循环学习网络MCCL,旨在:(1) 在帧级别检测多个概念,利用这些概念增强视频特征并提供时间事件线索;(2) 在描述网络内设计生成器与定位器之间的循环协同学习,促进语义感知与事件定位。具体地,对每帧进行弱监督概念检测,将检测到的概念嵌入融合进视频特征以提供事件提示;同时引入视频级概念对比学习,获得更具区分性的概念嵌入。在描述网络中,建立循环协同策略:生成器通过语义匹配引导定位器进行事件定位,定位器则通过位置匹配增强生成器的事件语义感知,使两者相互促进。MCCL在ActivityNet Captions和YouCook2数据集上达到当前最优性能。大量实验验证了其有效性和可解释性。
原文摘要 · Abstract (English)
Dense video captioning aims to detect and describe all events in untrimmed videos. This paper presents a dense video captioning network called Multi-Concept Cyclic Learning (MCCL), which aims to: (1) detect multiple concepts at the frame level, using these concepts to enhance video features and provide temporal event cues; and (2) design cyclic co-learning between the generator and the localizer within the captioning network to promote semantic perception and event localization. Specifically, we perform weakly supervised concept detection for each frame, and the detected concept embeddings are integrated into the video features to provide event cues. Additionally, video-level concept contrastive learning is introduced to obtain more discriminative concept embeddings. In the captioning network, we establish a cyclic co-learning strategy where the generator guides the localizer for event localization through semantic matching, while the localizer enhances the generator's event semantic perception through location matching, making semantic perception and event localization mutually beneficial. MCCL achieves state-of-the-art performance on the ActivityNet Captions and YouCook2 datasets. Extensive experiments demonstrate its effectiveness and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。