提升单模态表征,让音视频弱监督解析更准
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

- 用相似性迁移标注预训练数据,强化单模态理解
- 软约束并行优化单模态与多模态特征,定位更准
- 适合做音视频事件检测与弱监督学习的研究者
弱监督音视频视频解析(AVVP)旨在仅使用粗粒度标签识别和定位视频中的音频、视觉及音视频事件。现有方法主要沿两个方向:预训练伪标签生成器以提供细粒度跨模态语义引导,或改进AVVP模型架构以增强音视频融合。然而,由于音视频信号通常未对齐,准确的视频解析本质上依赖于对单模态事件的精确感知。现有方法过度强调多模态融合,忽视对单模态语义的引导与保留,导致伪标签噪声大、解析性能不佳。本文提出新框架,通过增强伪标签生成器与AVVP模型的单模态表征能力。具体地,引入基于相似性的标签迁移方法标注预训练数据,使生成器更好地理解单模态事件;同时采用软约束方式,在多模态融合的同时优化单模态特征建模。该设计实现单模态与跨模态表示的协同关注,显著提升事件定位性能。大量实验表明,本方法在伪标签生成与AVVP任务上均优于现有最先进方法。
原文摘要 · Abstract (English)
Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using only coarse-grained labels. Faced with the challenging task settings, existing research advances along two main paths: pre-training pseudo-label generators for fine-grained cross-modal semantic guidance, or refining AVVP model architectures to enhance audio-visual fusion. However, since audio and visual signals are typically unaligned, achieving accurate video parsing fundamentally relies on precise perception of uni-modal events. Yet these multi-modal focused strategies excessively emphasize multi-modal fusion while inadequately guiding and preserving uni-modal semantics, resulting in noisy pseudo-labels and sub-optimal video parsing performance. This paper proposes a novel framework that enhances uni-modal representations for both the pseudo-label generator and the AVVP model. Specifically, we introduce a similarity-based label migration approach to annotate pre-training data, thereby enabling the pseudo-label generator to better understand uni-modal events. We also employ a soft-constrained manner to refine modeling of uni-modal features in parallel with multi-modal fusion. These designs enable coordinated attention to both uni-modal and cross-modal representations, thus boosting the localization performance for events. Extensive experiments show that our method outperforms state-of-the-art methods in both pseudo-label and AVVP performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。