arXiv:2509.07341eess.AS2025-09被引 5

提出新模型同时降噪与补偿听力损失,提升助听器在嘈杂环境表现。

Affine Modulation-based Audiogram Fusion Network for Joint Noise Reduction and Hearing Loss Compensation

  • 用声学图谱和频谱特征融合,统一处理降噪与听力补偿
  • 在 HASQI、PESQ 等指标上超越当前最优联合模型
  • 适合开发更智能的个性化助听系统

助听器广泛用于为听力受损者提供个性化语音增强服务,改善生活质量。然而,在噪声环境中,助听器性能显著下降,因其将降噪(NR)与听力损失补偿(HLC)视为独立任务,导致缺乏系统性优化,忽略两者间交互关系,并增加系统复杂度。为此,我们提出一种新型声学图谱融合网络 AFN-HearNet,通过融合跨域声学图谱与频谱特征,同时解决 NR 与 HLC 任务。设计声学图谱专用编码器,将稀疏的声学图谱转换为深层表示,解决特征对齐问题。提出基于仿射调制的声学图谱融合频时 Conformer,自适应地将两类特征融合为统一深度表示以重构语音。此外,引入语音活动检测辅助训练任务,隐式将语音与非语音模式嵌入统一表示中。在多个数据集上的全面实验验证了各模块有效性。结果表明,AFN-HearNet 在关键指标如 HASQI 与 PESQ 上显著优于现有先进联合模型,实现了性能与效率的良好权衡。源代码与数据将发布于 https://github.com/deepnetni/AFN-HearNet。

原文摘要 · Abstract (English)

Hearing aids (HAs) are widely used to provide personalized speech enhancement (PSE) services, improving the quality of life for individuals with hearing loss. However, HA performance significantly declines in noisy environments as it treats noise reduction (NR) and hearing loss compensation (HLC) as separate tasks. This separation leads to a lack of systematic optimization, overlooking the interactions between these two critical tasks, and increases the system complexity. To address these challenges, we propose a novel audiogram fusion network, named AFN-HearNet, which simultaneously tackles the NR and HLC tasks by fusing cross-domain audiogram and spectrum features. We propose an audiogram-specific encoder that transforms the sparse audiogram profile into a deep representation, addressing the alignment problem of cross-domain features prior to fusion. To incorporate the interactions between NR and HLC tasks, we propose the affine modulation-based audiogram fusion frequency-temporal Conformer that adaptively fuses these two features into a unified deep representation for speech reconstruction. Furthermore, we introduce a voice activity detection auxiliary training task to embed speech and non-speech patterns into the unified deep representation implicitly. We conduct comprehensive experiments across multiple datasets to validate the effectiveness of each proposed module. The results indicate that the AFN-HearNet significantly outperforms state-of-the-art in-context fusion joint models regarding key metrics such as HASQI and PESQ, achieving a considerable trade-off between performance and efficiency. The source code and data will be released at https://github.com/deepnetni/AFN-HearNet.

助听器降噪听力补偿融合网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。