arXiv:2507.15396cs.SDcs.AI2025-07

轻量神经模型实现高保真听力损失仿真,实时性提升46倍。

Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation

  • 基于个性化听觉图编码的端到端模型,支持并行推理
  • 保持原MSBG的语音可懂度与感知质量,相关系数达0.92(STOI)
  • 1秒音频仿真时间从0.97秒降至0.021秒,适合嵌入式系统

听力损失仿真模型对助听器部署至关重要。然而现有模型计算复杂度高、延迟大,难以实现实时应用,且难以直接集成至语音处理系统。为此,我们提出Neuro-MSBG,一种轻量级端到端模型,采用个性化听觉图编码器实现高效的时频建模。实验表明,Neuro-MSBG支持并行推理,保留了原始MSBG的可懂度与感知质量,短时客观可懂度(STOI)的斯皮尔曼等级相关系数(SRCC)为0.9247,感知语音质量评估(PESQ)为0.8671。该模型将仿真耗时降低46倍(1秒输入从0.970秒降至0.021秒),显著提升效率与实用性。

原文摘要 · Abstract (English)

Hearing loss simulation models are essential for hearing aid deployment. However, existing models have high computational complexity and latency, which limits real-time applications and lack direct integration with speech processing systems. To address these issues, we propose Neuro-MSBG, a lightweight end-to-end model with a personalized audiogram encoder for effective time-frequency modeling. Experiments show that Neuro-MSBG supports parallel inference and retains the intelligibility and perceptual quality of the original MSBG, with a Spearman's rank correlation coefficient (SRCC) of 0.9247 for Short-Time Objective Intelligibility (STOI) and 0.8671 for Perceptual Evaluation of Speech Quality (PESQ). Neuro-MSBG reduces simulation runtime by a factor of 46 (from 0.970 seconds to 0.021 seconds for a 1 second input), further demonstrating its efficiency and practicality.

听力仿真端到端语音处理轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。