提出音频少样本学习新基准,揭示模型依赖背景噪声的致命弱点
SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

- 通过分离前景事件与背景环境,构建可控制的多层级上下文评估体系
- 多数先进模型在背景关联被破坏时性能骤降,连大模型也难幸免
- 揭示不同算法对虚假相关性的敏感差异,指导更鲁棒模型设计
少样本分类(FSC)常用于有限标注数据场景,但现有评估隐含假设目标概念与上下文无关。现实中,例子常伴随丰富上下文,使模型利用前景内容与背景信号间的虚假相关性。尽管图像领域已研究此问题,音频领域仍缺乏系统探索,现有基准对上下文结构控制有限。我们提出SpurAudio基准,利用音频中前景事件与背景环境的自然可分性,实现支持集与查询集间上下文变化的可控、多层级评估。实验表明,众多先进少样本方法在背景关联被破坏时性能严重下降,即使在大型预训练音频基础模型上亦然,排除了主干网络容量不足的解释。此外,传统基准表现相近的方法在虚假相关性敏感度上差异显著,揭示了特征表示与分类头在推理时交互带来的系统性优劣。这些发现为音频少样本学习提供了新洞见,强调评估需显式考察上下文依赖性。
原文摘要 · Abstract (English)
Few-shot classification (FSC) is widely used for learning from limited labeled data, yet most evaluations implicitly assume that target concepts are independent of contextual cues. In real-world settings, however, examples often appear within rich contexts, allowing models to exploit spurious correlations between foreground content and background signals. While such effects have been studied in few-shot image classification, their role in few-shot audio classification remains largely unexplored, and existing audio benchmarks offer limited control over contextual structure. We introduce SpurAudio, a benchmark that leverages the natural separability of foreground events and background environments in audio to enable controlled, multi-level evaluation of contextual shifts across support and query sets. Using this benchmark, we show that many state-of-the-art few-shot methods suffer severe performance degradation when background correlations are disrupted, despite achieving similar accuracy under standard evaluation protocols. Crucially, this vulnerability persists even in large pretrained audio foundation models, ruling out limited backbone capacity as an explanation. Moreover, methods that appear comparable under conventional benchmarks can exhibit markedly different sensitivity to spurious correlations, revealing systematic algorithmic strengths and vulnerabilities tied to how feature representations interact with classifier heads at inference time. These findings provide new insight into the behavior of few-shot methods in audio and highlight the need for benchmarks that explicitly probe context dependence when evaluating FSC models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。