提出新模型实现长视频中音视频事件的精准定位
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
- 通过跨模态一致性与多时序粒度协作机制融合音视频信息
- 在UnAV-100数据集上达到当前最优性能,显著提升重叠事件识别率
- 适合关注长视频理解、多模态事件检测的研究者
在音视频学习领域,多数研究集中于短视频任务。本文聚焦更具实际意义的密集音视频事件定位(DAVEL)任务,推动长时未剪辑视频中的音视频场景理解。该任务旨在同时识别并精确标注音频与视觉流中所有并发事件。每段视频包含多种类别、时间重叠且持续时间各异的密集事件。为此,有效利用音视频关联性及多粒度时序特征至关重要。本文提出新型CCNet,包含两个核心模块:跨模态一致性协作(CMCC)与多时序粒度协作(MTGC)。CMCC模块含交叉模态交互分支与时间一致性门控分支:前者通过编码音视频关系聚合一致事件语义,后者引导一模态聚焦另一模态识别的关键事件时区。MTGC模块包含粗到细与细到粗协作块,实现粗粒度与细粒度时序特征的双向支持。在UnAV-100数据集上的大量实验验证了模块设计的有效性,实现了密集音视频事件定位的新基准性能。代码已开源。
原文摘要 · Abstract (English)
In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for longer, untrimmed videos. This task seeks to identify and temporally pinpoint all events simultaneously occurring in both audio and visual streams. Typically, each video encompasses dense events of multiple classes, which may overlap on the timeline, each exhibiting varied durations. Given these challenges, effectively exploiting the audio-visual relations and the temporal features encoded at various granularities becomes crucial. To address these challenges, we introduce a novel CCNet, comprising two core modules: the Cross-Modal Consistency Collaboration (CMCC) and the Multi-Temporal Granularity Collaboration (MTGC). Specifically, the CMCC module contains two branches: a cross-modal interaction branch and a temporal consistency-gated branch. The former branch facilitates the aggregation of consistent event semantics across modalities through the encoding of audio-visual relations, while the latter branch guides one modality's focus to pivotal event-relevant temporal areas as discerned in the other modality. The MTGC module includes a coarse-to-fine collaboration block and a fine-to-coarse collaboration block, providing bidirectional support among coarse- and fine-grained temporal features. Extensive experiments on the UnAV-100 dataset validate our module design, resulting in a new state-of-the-art performance in dense audio-visual event localization. The code is available at https://github.com/zzhhfut/CCNet-AAAI2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。