首次构建生成式零样本声音分类基准,验证生成模型有效性。
A Benchmark of Generative Methods for Zero-Shot Environmental Sound Classification
- 对比四种生成方法:变分、对抗、扩散与去噪网络
- CGDN表现最佳,平均准确率超越其他生成模型
- 强调优化稳定性对生成式零样本学习至关重要
零样本学习使模型能利用语义信息泛化到未见类别,弥合训练类与未知测试类之间的差距。尽管在计算机视觉中研究广泛,其在环境音频领域的应用仍较薄弱,生成方法更是少有关注。本文首次构建了生成式零样本环境声音分类的基准。评估了四种范式:变分、对抗、扩散和去噪,包括来自计算机视觉的CADA-VAE和LisGAN,以及本文提出的两种嵌入生成方法:基于去噪扩散概率模型(DDPM)和条件生成去噪网络(CGDN)。在五个环境音频数据集(ESC-50、ARCA23K-FSD、FSC22、UrbanSound8K、TAU Urban Acoustic Scenes 2019)和一个音乐数据集(GTZAN)上的实验表明,生成方法可与现有基于兼容性的方法媲美。其中,CGDN取得最高平均准确率,显著优于基于DDPM和GAN的方法,且与强基线ALE无统计差异。结果表明,优化稳定性是生成式零样本学习中的关键因素。
原文摘要 · Abstract (English)
Zero-shot learning enables models to generalise to unseen classes using semantic information, bridging the gap between training classes and previously unseen test classes. While widely studied in computer vision, its application to environmental audio remains underexplored, and generative approaches have received little attention. This work presents the first benchmark of generative methods for zero-shot environmental sound classification. Four approaches spanning variational, adversarial, diffusion-based, and denoising paradigms are evaluated. The benchmark includes CADA-VAE and LisGAN, adapted from computer vision, together with two embedding-generation methods introduced in this work: one based on a denoising diffusion probabilistic model (DDPM) and the other on a conditional generative denoising network (CGDN). Experiments on five environmental audio datasets (ESC-50, ARCA23K-FSD, FSC22, UrbanSound8K, and TAU Urban Acoustic Scenes 2019) and one music dataset (GTZAN) show that generative methods are competitive with established compatibility-based approaches. Among the evaluated generative methods, CGDN achieves the highest average accuracy and is the only one to significantly outperform both the DDPM- and GAN-based methods, while remaining statistically indistinguishable from the strong ALE baseline. These findings suggest that optimisation stability is an important factor in generative zero-shot learning for environmental audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。