arXiv:2607.25887eess.AScs.AI2026-07

对比两种方法在设备差异下的音景分类表现,发现DANN更稳定可靠。

Device Invariance using Domain Adaptation on Acoustic Scene Classification

  • 用DANN和CDAN处理不同设备录音的特征差异
  • DANN对CNN与Transformer都有效,CDAN仅适配CNN
  • 适合关注跨设备音频识别的研究者

本文研究了在使用基于卷积神经网络(CNN)和基于Transformer的特征表示进行音景分类时,领域自适应技术的有效性。评估了两种知名领域自适应方法:领域对抗神经网络(DANN)和条件领域对抗网络(CDAN),在多种领域偏移情况下的表现。实验表明,DANN对两种特征提取器均能提供较稳定的领域自适应效果;而CDAN仅在基于CNN的特征提取器上表现有效。该研究揭示了领域自适应方法需根据底层特征表示进行适配。在DCASE 2020数据集上,通过多个设备的实证验证了上述结论。

原文摘要 · Abstract (English)

This paper explores the effectiveness of domain adaptation techniques when using convolutional neural network (CNN)-based and transformer-based feature representations for acoustic scene classification. Two well-known domain adaptation techniques, namely domain adversarial neural network (also called DANN) and conditional domain adversarial network (also called CDAN) are evaluated under various domain shifts. Our study indicates that DANN provides effective domain adaptation fairly consistently for both feature extractors. On the other hand, CDAN provides effective domain adaptation only for CNN-based feature extractors. The study gives insights into how domain adaptation methods may need to be tailored to the underlying feature representation. Experimental evaluation with multiple devices on the DCASE 2020 dataset supports the observations.

音景分类领域自适应CNNTransformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。