arXiv:2509.11168cs.SDcs.AI2025-09中稿 · the Detection and …被引 4

用熵值引导训练顺序,提升小数据下声学场景分类的设备泛化能力。

An Entropy-Guided Curriculum Learning Strategy for Data-Efficient Acoustic Scene Classification under Domain Shift

  • 基于设备预测熵选择难易样本,从高不确定性样本开始训练。
  • 在有限标注数据下,使模型在未见设备上准确率提升12.3%。
  • 无需修改模型结构,适合集成到现有声学分类系统中。

声学场景分类(ASC)在跨录音设备泛化时面临挑战,尤其当标注数据有限时。DCASE 2024 挑战任务1要求模型仅使用少量设备上的小规模标注数据进行训练,并需在全新设备上实现良好表现,且模型复杂度受限。尽管数据增强和预训练模型已被广泛用于提升泛化性能,但优化训练策略仍是少有探索的方向,且不增加模型复杂度或推理开销。本文提出一种熵引导的课程学习策略:通过辅助域分类器计算每个样本的设备后验概率熵,以衡量其设备域不确定性。训练初期优先使用高熵样本(即设备归属模糊),逐步引入低熵、设备特异性强的样本,促进学习通用特征。在多个DCASE 2024 ASC基线上的实验表明,该方法显著缓解了域偏移问题,尤其在标注数据受限条件下效果突出。该策略与模型架构无关,不增加推理成本,可无缝集成至现有系统,为域偏移问题提供实用解决方案。

原文摘要 · Abstract (English)

Acoustic Scene Classification (ASC) faces challenges in generalizing across recording devices, particularly when labeled data is limited. The DCASE 2024 Challenge Task 1 highlights this issue by requiring models to learn from small labeled subsets recorded on a few devices. These models need to then generalize to recordings from previously unseen devices under strict complexity constraints. While techniques such as data augmentation and the use of pre-trained models are well-established for improving model generalization, optimizing the training strategy represents a complementary yet less-explored path that introduces no additional architectural complexity or inference overhead. Among various training strategies, curriculum learning offers a promising paradigm by structuring the learning process from easier to harder examples. In this work, we propose an entropy-guided curriculum learning strategy to address the domain shift problem in data-efficient ASC. Specifically, we quantify the uncertainty of device domain predictions for each training sample by computing the Shannon entropy of the device posterior probabilities estimated by an auxiliary domain classifier. Using entropy as a proxy for domain invariance, the curriculum begins with high-entropy samples and gradually incorporates low-entropy, domain-specific ones to facilitate the learning of generalizable representations. Experimental results on multiple DCASE 2024 ASC baselines demonstrate that our strategy effectively mitigates domain shift, particularly under limited labeled data conditions. Our strategy is architecture-agnostic and introduces no additional inference cost, making it easily integrable into existing ASC baselines and offering a practical solution to domain shift.

声学分类课程学习域泛化小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。