arXiv:2509.09262cs.SDcs.AI2025-09被引 1

用设备感知的教师模型,让小模型更好识别不同设备的声音场景。

Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification

  • 用双教师蒸馏框架,一个通用、一个专注设备鲁棒性的教师协同教学。
  • 在未见设备上准确率达57.93%,显著优于基线。
  • 适合资源受限设备部署,且能利用测试时的设备标签提升性能。

本文介绍我们针对DCASE 2025挑战赛任务1——低复杂度设备鲁棒声景分类的提交方案。系统同时应对严格复杂度约束与对已见及未见设备的泛化需求,并利用允许在测试时使用设备标签的新规则。所提方法基于知识蒸馏框架,由轻量级CP-MobileNet学生模型从一个紧凑的双教师集成中学习。该集成包含一个使用标准交叉熵训练的基准PaSST教师,以及一个采用新提出的设备感知特征对齐(DAFA)损失训练的‘泛化专家’教师,后者显式构建利于设备鲁棒性的特征空间。为充分利用测试时的设备标签,学生模型在蒸馏后还经历一次设备特定微调阶段。最终系统在开发集上达到57.93%的准确率,尤其在未见设备上表现显著优于官方基线。

原文摘要 · Abstract (English)

In this technical report, we describe our submission for Task 1, Low-Complexity Device-Robust Acoustic Scene Classification, of the DCASE 2025 Challenge. Our work tackles the dual challenges of strict complexity constraints and robust generalization to both seen and unseen devices, while also leveraging the new rule allowing the use of device labels at test time. Our proposed system is based on a knowledge distillation framework where an efficient CP-MobileNet student learns from a compact, specialized two-teacher ensemble. This ensemble combines a baseline PaSST teacher, trained with standard cross-entropy, and a 'generalization expert' teacher. This expert is trained using our novel Device-Aware Feature Alignment (DAFA) loss, adapted from prior work, which explicitly structures the feature space for device robustness. To capitalize on the availability of test-time device labels, the distilled student model then undergoes a final device-specific fine-tuning stage. Our proposed system achieves a final accuracy of 57.93\% on the development set, demonstrating a significant improvement over the official baseline, particularly on unseen devices.

知识蒸馏声景分类设备鲁棒轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。