arXiv:2501.04116eess.AScs.SD2025-01被引 5

提出无伪影神经网络dCoNNear,提升听觉辅助与语音增强音质。

dCoNNear: An Artifact-Free Neural Network Architecture for Closed-loop Audio Signal Processing

  • 设计新架构dCoNNear,避免非理想采样引入的谐波与混叠伪影。
  • 在听觉模型与语音增强任务中均实现音质提升,无新伪影产生。
  • 适合高保真音频应用,尤其对个性化助听器开发有重要价值。

深度神经网络(DNN)显著提升了语音增强、合成及助听器算法等音频处理应用。基于DNN的闭环系统因性能鲁棒且能适应多样环境而广受青睐。然而,现有方法常因次优采样方式导致音质退化与伪影出现。本文提出dCoNNear,一种专为闭环框架设计的新DNN架构,旨在消除由非理想采样层引发的典型伪影——尤其是谐波与混叠伪影。通过一个原理验证案例,在包含正常与听力受损者生物物理听觉模型的闭环框架中,成功构建个性化助听器算法。进一步在语音增强实验中验证其广泛适用性,结果表明dCoNNear不仅能准确模拟传统非DNN生物物理模型的所有处理阶段,还能在助听器与语音增强任务中有效消除可听伪影,显著提升感知音质。本研究提供了一个稳健、感知透明的闭环音频处理框架,适用于高保真音频应用。

原文摘要 · Abstract (English)

Recent advances in deep neural networks (DNNs) have significantly improved various audio processing applications, including speech enhancement, synthesis, and hearing-aid algorithms. DNN-based closed-loop systems have gained popularity in these applications due to their robust performance and ability to adapt to diverse conditions. Despite their effectiveness, current DNN-based closed-loop systems often suffer from sound quality degradation caused by artifacts introduced by suboptimal sampling methods. To address this challenge, we introduce dCoNNear, a novel DNN architecture designed for seamless integration into closed-loop frameworks. This architecture specifically aims to prevent the generation of spurious artifacts-most notably tonal and aliasing artifacts arising from non-ideal sampling layers. We demonstrate the effectiveness of dCoNNear through a proof-of-principle example within a closed-loop framework that employs biophysically realistic models of auditory processing for both normal and hearing-impaired profiles to design personalized hearing-aid algorithms. We further validate the broader applicability and artifact-free performance of dCoNNear through speech-enhancement experiments, confirming its ability to improve perceptual sound quality without introducing architecture-induced artifacts. Our results show that dCoNNear not only accurately simulates all processing stages of existing non-DNN biophysical models but also significantly improves sound quality by eliminating audible artifacts in both hearing-aid and speech-enhancement applications. This study offers a robust, perceptually transparent closed-loop processing framework for high-fidelity audio applications.

音频处理神经网络助听器音质优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。