arXiv:2607.04826eess.AScs.SD2026-07中稿 · ICASSP 2026

小模型专注说话人身份,语音增强效果远超大模型。

Ranking the Impact of Contextual Specialization in Neural Speech Enhancement

  • 用说话人身份等上下文信息微调小模型,提升语音质量。
  • 专用于特定说话人和噪声的小模型超越十倍大的通用模型。
  • 语言特征对性能提升关键,适合助听器等低资源场景。

我们系统研究了从约10千参数到中大型(约200万至500万参数)的神经语音增强系统,这些系统利用说话人身份、噪声类型、性别、语言及信噪比(SNR)等上下文信息进行针对性优化。通过在特定数据子集上微调通用模型,发现专注于说话人身份能持续带来最大语音可懂度与质量提升;而针对SNR、噪声类型或性别则收益甚微。关键发现:一个专用于特定说话人和噪声类型的小微模型,其性能可媲美甚至超过规模为其十倍的通用模型。跨语言测试表明,针对目标语言微调的模型优于多语言通用模型,说明语言是重要的可专化特征。该研究揭示了小型自适应模型在助听器等资源受限场景中的巨大潜力。

原文摘要 · Abstract (English)

We systematically investigate neural speech enhancement systems, ranging from very small ($\sim$10\,k parameters) to medium-large ($\sim$2-5\,M parameters), which specialize to acoustic conditions using contextual information such as speaker identity, noise type, speaker gender, spoken language, and SNR. By fine-tuning generalist models on specific data subsets, we find that specializing to a speaker's identity consistently yields the largest gains in estimated speech intelligibility and quality. In contrast, specializing to SNR, noise type, or gender offers only marginal benefits. Crucially, we show that a small model specialized to both a specific speaker and a specific noise type can match or exceed the performance of a generalist model ten times its size. Further, cross-lingual tests reveal that models specialized to a target language outperform multilingual generalists, suggesting that language is a salient feature for specialization. These findings highlight the potential of small, adaptive models for resource-constrained applications like hearing aids, which specialize on-the-fly to contextual information.

语音增强小模型上下文专用说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。