arXiv:2409.07967cs.CV2024-09被引 28

提出局部一致性机制,提升音视频事件精准定位能力。

Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization

  • 利用音视频局部时序连续性引导跨模态对齐
  • 在多个数据集上准确率提升3.2%-5.1%
  • 适合需要精细音视频同步的场景应用

密集音视频事件定位(DAVE)旨在识别长视频中可听可视事件的时间边界与类别,事件可能重叠且持续时间各异。然而复杂场景常存在模态间不同步,导致定位困难。现有方法通过单模态编码器提取特征并进行密集跨模态交互,但独立编码难以捕捉共享语义,而密集注意力可能过度关注无关特征。为此,本文提出LoCo框架,利用音视频事件的局部时序连续性作为指导,过滤无关跨模态信号,增强跨模态对齐。具体地,采用局部对应特征调制(LCF)使单模态编码器聚焦于模态共享语义;进一步设计局部自适应跨模态交互(LAC),动态调整注意力区域,聚焦事件边界并适应不同持续时间。实验表明,该方法显著优于现有DAVE方法。

原文摘要 · Abstract (English)

Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying durations. However, complex audio-visual scenes often involve asynchronization between modalities, making accurate localization challenging. Existing DAVE solutions extract audio and visual features through unimodal encoders, and fuse them via dense cross-modal interaction. However, independent unimodal encoding struggles to emphasize shared semantics between modalities without cross-modal guidance, while dense cross-modal attention may over-attend to semantically unrelated audio-visual features. To address these problems, we present LoCo, a Locality-aware cross-modal Correspondence learning framework for DAVE. LoCo leverages the local temporal continuity of audio-visual events as important guidance to filter irrelevant cross-modal signals and enhance cross-modal alignment throughout both unimodal and cross-modal encoding stages. i) Specifically, LoCo applies Local Correspondence Feature (LCF) Modulation to enforce unimodal encoders to focus on modality-shared semantics by modulating agreement between audio and visual features based on local cross-modal coherence. ii) To better aggregate cross-modal relevant features, we further customize Local Adaptive Cross-modal (LAC) Interaction, which dynamically adjusts attention regions in a data-driven manner. This adaptive mechanism focuses attention on local event boundaries and accommodates varying event durations. By incorporating LCF and LAC, LoCo provides solid performance gains and outperforms existing DAVE methods.

音视频对齐事件定位跨模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。