arXiv:2603.07471eess.AScs.AI2026-03中稿 · ICASSP 2026

轻量级语音增强模型在线自适应,仅更新1%参数即可提升音质。

Towards Lightweight Adaptation of Speech Enhancement Models in Real-World Environments

  • 用低秩适配器在冻结主干上实现轻量自适应
  • 每场景仅20次更新,平均提升1.51 dB SI-SDR
  • 适合真实环境中资源受限设备部署

近期研究表明,部署后自适应可提升语音增强模型在未知噪声条件下的鲁棒性。然而,现有方法常带来高昂的计算与内存开销,限制其在设备端的应用。本文研究动态声学场景变化下的真实环境适应问题,提出一种轻量级框架:在冻结主干基础上添加低秩适配器,并通过自监督训练更新。在覆盖37种噪声类型、111个环境、信噪比范围涵盖[-8, 0] dB的连续场景评估中,该方法更新参数少于基线模型的1%,仅需每场景20次更新即实现平均1.51 dB的SI-SDR提升。相比当前最优方法,本框架在感知质量上表现相当或更优,收敛更平稳,验证了其在真实声学条件下设备端轻量自适应的实用性。

原文摘要 · Abstract (English)

Recent studies have shown that post-deployment adaptation can improve the robustness of speech enhancement models in unseen noise conditions. However, existing methods often incur prohibitive computational and memory costs, limiting their suitability for on-device deployment. In this work, we investigate model adaptation in realistic settings with dynamic acoustic scene changes and propose a lightweight framework that augments a frozen backbone with low-rank adapters updated via self-supervised training. Experiments on sequential scene evaluations spanning 111 environments across 37 noise types and three signal-to-noise ratio ranges, including the challenging [-8, 0] dB range, show that our method updates fewer than 1% of the base model's parameters while achieving an average 1.51 dB SI-SDR improvement within only 20 updates per scene. Compared to state-of-the-art approaches, our framework achieves competitive or superior perceptual quality with smoother and more stable convergence, demonstrating its practicality for lightweight on-device adaptation of speech enhancement models under real-world acoustic conditions.

语音增强轻量级适配自监督学习设备端部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。