arXiv:2506.11331cs.AIcs.SD2025-06

轻量级无监督域适应框架,让低功耗设备精准识别城市混响声音。

MUDAS: Mote-scale Unsupervised Domain Adaptation in Multi-label Sound Classification

  • 仅重训分类器,用高置信度数据本地适配模型。
  • 在纽约多地点数据上准确率显著优于现有方法。
  • 适合边缘设备部署,支持复杂声学环境下的多标签识别。

无监督域适应(UDA)对适应新环境中的机器学习模型至关重要,因数据分布偏移常导致性能下降。现有UDA算法针对单标签任务设计,依赖大量计算资源,难以应用于多标签场景及资源受限的物联网(IoT)设备。在城市声音分类中,重叠声源与变化声学特性要求低功耗设备具备强健的多标签自适应能力。为此,本文提出面向声音的极小尺度无监督域适应框架(MUDAS),专为资源受限的IoT环境中的多标签声音分类设计。MUDAS通过选择性地在设备端使用高置信度数据重新训练分类器,大幅降低计算与内存开销,实现就地模型适配。同时,引入类别自适应阈值生成可靠伪标签,并采用多样性正则化提升多标签分类精度。在纽约多个地点采集的SONYC-UST数据集上的实验表明,MUDAS在资源受限条件下显著优于现有UDA算法,展现出良好性能。

原文摘要 · Abstract (English)

Unsupervised Domain Adaptation (UDA) is essential for adapting machine learning models to new, unlabeled environments where data distribution shifts can degrade performance. Existing UDA algorithms are designed for single-label tasks and rely on significant computational resources, limiting their use in multi-label scenarios and in resource-constrained IoT devices. Overcoming these limitations is particularly challenging in contexts such as urban sound classification, where overlapping sounds and varying acoustics require robust, adaptive multi-label capabilities on low-power, on-device systems. To address these limitations, we introduce Mote-scale Unsupervised Domain Adaptation for Sounds (MUDAS), a UDA framework developed for multi-label sound classification in resource-constrained IoT settings. MUDAS efficiently adapts models by selectively retraining the classifier in situ using high-confidence data, minimizing computational and memory requirements to suit on-device deployment. Additionally, MUDAS incorporates class-specific adaptive thresholds to generate reliable pseudo-labels and applies diversity regularization to improve multi-label classification accuracy. In evaluations on the SONYC Urban Sound Tagging (SONYC-UST) dataset recorded at various New York City locations, MUDAS demonstrates notable improvements in classification accuracy over existing UDA algorithms, achieving good performance in a resource-constrained IoT setting.

声音分类边缘计算无监督学习多标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。