用深度神经网络统一估计语音、音乐、噪声的频时功率,提升助听器环境识别效率。
DNN-Based Frequency-Dependent Estimation of Speech, Music, and Noise Power in Acoustic Mixtures for Hearing-Aid Scene Analysis

- 设计因果低复杂度DNN,联合估计语音/音乐/噪声在频时域的功率分布
- 基于该表示的语音活动检测性能媲美顶尖方法,且可衍生多种分析任务
- 适合需要多任务协同处理的实时助听设备场景分析应用
声学场景分析对自适应助听器信号处理至关重要。现有先进系统通常依赖多个独立估算器完成场景分类、语音活动检测(VAD)或信噪比估计等任务,导致计算开销大且未能利用相关任务间的内在关联。为解决此问题,我们提出一种统一且可解释的声学场景表征方法,将观测混合信号谱分解为语音、音乐和噪声的功率分量,这一思路契合助听用户典型的听觉关注目标。具体而言,采用因果低复杂度深度神经网络,估计时频依赖的功率比例,从而可通过简单后处理推导出多种下游声学场景分析指标。本文以语音活动检测(VAD)为例验证该表示的有效性,结果表明其性能可与最先进估计算法相当,同时提供更丰富的场景描述信息。
原文摘要 · Abstract (English)
Acoustic scene analysis is essential for adapting hearing-aid signal processing algorithms to the current listening environment. However, state-of-the-art (SOTA) systems typically rely on multiple independent estimators for tasks such as scene classification, Voice Activity Detection (VAD), or Signal-to-Noise Ratio estimation, which increases computational complexity and fails to exploit dependencies between related tasks. To address this problem, we propose a unified and interpretable acoustic scene representation by decomposing the observed mixture spectrum into speech, music, and noise power components. This is motivated by the typical listening targets of hearing-aid users. In particular, we estimate time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing. In this work, we validate the proposed representation using VAD as a representative downstream task and show performance comparable to a SOTA estimator while providing a substantially richer scene description.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。