arXiv:2605.17512eess.AScs.SD2026-05

针对音频标签数据中类别级标注不可靠问题,提出按类别调控监督强度的新方法。

Robust Audio Tagging under Class-wise Supervision Unreliability

论文配图:Robust Audio Tagging under Class-wise Supervision Unreliability
图 1 · 摘自论文原文
  • 为每个声音类别学习独立的不可靠参数,动态降低不可靠标签权重
  • 在真实与生成音频混合的数据上提升模型鲁棒性,准确率提升显著
  • 适用于弱监督音频识别场景,尤其适合标注质量不一的大规模数据

弱标签数据集如 AudioSet 推动了音频标记的发展,但不同声音类别的标注质量差异显著。标签可能存在缺失、模糊或不可靠,导致优化过程中产生类别依赖的监督偏差。当真实与生成音频在训练中日益混合时,该问题更严重,因为生成样本未必匹配其语义标签。以往工作主要关注缺失正例标签的问题,而本文聚焦于三类新来源的不可靠监督:虚假添加、相似类别误分配、标签证据减弱。这些影响引入了未被现有方法显式建模的类别依赖优化偏差。为此,本文提出类别级监督不可靠性(CSU)框架,在训练中按类别控制监督强度。CSU 为每个类别学习独立的不可靠参数,通过降权不可靠监督来改进模型,无需修改模型结构或推理流程。为支持评估,还构建了 ESC-FreeGen50——一个包含 50 个声音类别的手动验证基准,融合真实与生成音频。在可控基准和 AudioSet 上的实验表明,CSU 在多种架构和不同不可靠监督源下均提升了鲁棒性。结果表明,显式建模类别级监督不可靠性是一种有效且实用的策略,适用于大规模弱监督音频标记。

原文摘要 · Abstract (English)

Weakly labeled datasets such as AudioSet have driven recent progress in audio tagging. However, annotation quality varies across sound classes. Labels may be incomplete, ambiguous, or unreliable, which introduces class-dependent supervision bias during optimisation. The issue becomes harder as real and generated audio are increasingly mixed in training, and generated samples do not always match their intended semantic labels. Prior work mainly addressed unreliable supervision from missing-positive labels, while this paper targets three other sources of unreliable supervision: spurious additions, misassignments between similar classes, and weakened label evidence. These effects introduce class-dependent optimisation bias that is not explicitly modeled by most existing methods. To bridge this gap, the paper proposes a Class-wise Supervision Unreliability (CSU) framework that controls supervision strength at the class level during training. CSU learns a separate unreliability parameter for each class and down-weights less reliable supervision without changing the model architecture or inference process. To support evaluations, this paper also introduces ESC-FreeGen50, a manually verified benchmark of 50 sound classes that combines real and generated audio. Experiments on controlled benchmarks and AudioSet show that CSU improves robustness across different architectures and different sources of supervision unreliability. The results indicate that explicit class-wise modeling of supervision unreliability is an effective and practical strategy for robust audio tagging under large-scale weakly labeled training. Code and data are available at: https://github.com/Yuanbo2020/CSU

音频识别弱监督鲁棒性生成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。