用深度学习实现个性化助听器端到端放大,效果显著优于传统方法。
NeuroAMP: A Novel End-to-end General Purpose Deep Neural Amplifier for Personalized Hearing Aids
- 输入频谱和听力图,用Transformer等模型端到端优化放大
- 在多个数据集上达到0.99以上相关性,音乐场景下仍保持高性能
- 新增降噪功能,比传统方法提升10%语音感知与质量得分
助听器普及率上升,但传统多模块集成的放大优化仍具挑战。本文提出NeuroAMP,一种基于深度神经网络的端到端个性化放大方案,输入包括频谱特征与用户听力图,测试了CNN、LSTM、CRNN和Transformer四种架构。进一步提出Denoising NeuroAMP,在放大基础上集成降噪能力。训练时采用覆盖TIMIT、TMHINT(语音)和Cadenza Challenge MUSIC(音乐)的综合数据增强策略。评估使用HASPI、HASQI和HAAQI指标,Transformer架构在TIMIT上分别达0.9927(HASQI)和0.9905(HASPI),在Cadenza Challenge MUSIC上达0.9738(HAAQI)。数据增强使模型在未见数据集(如VCTK、MUSDB18-HQ)仍表现稳定。Denoising NeuroAMP在VoiceBank+DEMAND上优于传统NAL-R+WDRC及两阶段基线,HASPI提升至0.90,HASQI达0.59,性能提高10%。结果表明,该方法可显著提升个性化助听器表现。
原文摘要 · Abstract (English)
The prevalence of hearing aids is increasing. However, optimizing the amplification processes of hearing aids remains challenging due to the complexity of integrating multiple modular components in traditional methods. To address this challenge, we present NeuroAMP, a novel deep neural network designed for end-to-end, personalized amplification in hearing aids. NeuroAMP leverages both spectral features and the listener's audiogram as inputs, and we investigate four architectures: Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), Convolutional Recurrent Neural Network (CRNN), and Transformer. We also introduce Denoising NeuroAMP, an extension that integrates noise reduction along with amplification capabilities for improved performance in real-world scenarios. To enhance generalization, a comprehensive data augmentation strategy was employed during training on diverse speech (TIMIT and TMHINT) and music (Cadenza Challenge MUSIC) datasets. Evaluation using the Hearing Aid Speech Perception Index (HASPI), Hearing Aid Speech Quality Index (HASQI), and Hearing Aid Audio Quality Index (HAAQI) demonstrates that the Transformer architecture within NeuroAMP achieves the best performance, with SRCC scores of 0.9927 (HASQI) and 0.9905 (HASPI) on TIMIT, and 0.9738 (HAAQI) on the Cadenza Challenge MUSIC dataset. Notably, our data augmentation strategy maintains high performance on unseen datasets (e.g., VCTK, MUSDB18-HQ). Furthermore, Denoising NeuroAMP outperforms both the conventional NAL-R+WDRC approach and a two-stage baseline on the VoiceBank+DEMAND dataset, achieving a 10% improvement in both HASPI (0.90) and HASQI (0.59) scores. These results highlight the potential of NeuroAMP and Denoising NeuroAMP to deliver notable improvements in personalized hearing aid amplification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。