用可微分听觉模型实现听力辅具个性化,提升嘈杂环境下的语音清晰度。
The Differentiable Auditory Loop (DAL): An ML Framework for Hyper-Personalized Hearing Aids

- 通过可微分耳蜗模型匹配听力受损与正常人的神经反应模式
- 优化后的网络在信号保真度和神经表征上优于现有助听器基线
- 适合需要高精度听力补偿的临床研究与个性化助听设备开发
传统助听器依赖固定频率增益与压缩处理,难以应对多说话人等复杂听觉环境(即‘鸡尾酒会问题’)。为更全面解决听力损失的编码功能障碍,本文提出可微分听觉回路(DAL),一个用于个性化助听器设计与适配的开源框架。首个实现中引入了CARFAC——一个可微分的人类耳蜗模型,移植至JAX,并用于优化深度神经网络,使其输出匹配受损听觉神经活动模式与正常听力参考。为实现精细的时频信号处理,采用SEANet(一种全卷积波形到波形的UNet生成器)进行建模。通过比较正常听力与个体化听力损伤的CARFAC输出,利用基于神经活动模式(NAP)与稳定听觉图像(SAI)的损失函数进行微调。梯度下降使SEANet同时完成输入去噪与听力损失补偿。在神经表征与信号保真度指标上,DAL优化的SEANet模型均优于测试的主助听器(MHA)基线。该框架为基于模型的、机器学习驱动的助听器个性化提供了可行路径。下一步将推进硬件部署以支持真实世界临床验证。
原文摘要 · Abstract (English)
Conventional hearing aids rely on fixed, frequency-dependent amplification and compression to manage reduced sensitivity, which often fails to provide sufficient listening support in complex environments, such as situations with multiple speakers (the ``cocktail party'' problem). To more comprehensively address the underlying encoding dysfunctions of hearing loss, we introduce the Differentiable Auditory Loop (DAL), a new open-source framework for personalized hearing aid design and fitting. Our first implementation of DAL incorporates CARFAC, a differentiable model of human cochlear function, which we ported to JAX, to optimize a deep neural network to match impaired auditory neural activity patterns with a normal-hearing reference. To build a hearing aid with the fine-grained spectro-temporal signal processing required, we adopt SEANet, a waveform-to-waveform fully convolutional UNet generator. We fine-tune the network by comparing the outputs of a CARFAC model fitted to normal hearing with that of a CARFAC model fitted to match each subject's individual hearing impairment. The comparison is done using loss functions derived from the respective CARFAC neural activity pattern (NAP) outputs and stabilized auditory images (SAIs), the latter providing a 2D representation that captures phase-insensitive temporal structure in the auditory nerve output. Through gradient descent, the SEANet model learns to both denoise the input and compensate for the hearing loss modelled by the impaired CARFAC model. Across neural-representation and signal-fidelity metrics, the DAL-optimized SEANet model outperformed the tested master hearing aid (MHA) baselines. The DAL framework provides a practical path toward model-based, machine-learning-driven personalization of hearing aid signal processing. Next steps include hardware deployment to enable real-world clinical testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。