arXiv:2509.14097cs.CVcs.MM2025-09被引 2

通过伪监督与跨模态对齐,提升弱监督音视频事件解析的精度。

Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing

  • 用EMA动态生成可靠片段级标签,替代仅靠视频级标签的粗粒度监督。
  • 在LLP和UnAV-100上达到最新最好结果,显著提升事件检测准确率。
  • 适合做音视频多模态理解、弱监督学习的研究者参考。

弱监督音视频视频解析(AVVP)旨在无需时间标注的情况下检测可听、可见及视听事件。以往方法主要依赖对比或协同学习优化全局预测,却忽视了稳定的片段级监督与类别感知的跨模态对齐。为此,我们提出两种策略:(1) 基于指数移动平均(EMA)的伪监督框架,通过自适应阈值或top-k选择生成可靠的片段级掩码,提供超越视频级标签的稳定时序指导;(2) 类别感知跨模态一致性(CMA)损失,对可靠片段-类别配对中的音视频嵌入进行对齐,确保模态间一致性并保留时序结构。在LLP与UnAV-100数据集上的评估显示,该方法在多个指标上达到当前最优性能。

原文摘要 · Abstract (English)

Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or collaborative learning, but neglected stable segment-level supervision and class-aware cross-modal alignment. To address this, we propose two strategies: (1) an exponential moving average (EMA)-guided pseudo supervision framework that generates reliable segment-level masks via adaptive thresholds or top-k selection, offering stable temporal guidance beyond video-level labels; and (2) a class-aware cross-modal agreement (CMA) loss that aligns audio and visual embeddings at reliable segment-class pairs, ensuring consistency across modalities while preserving temporal structure. Evaluations on LLP and UnAV-100 datasets shows that our method achieves state-of-the-art (SOTA) performance across multiple metrics.

音视频解析弱监督跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。