用动态声源方向引导深度波束成形,提升语音增强的空间一致性。
Interpretable Binaural Deep Beamforming Guided by Time-Varying Relative Transfer Function
- 基于时变相对传递函数追踪声源方向,指导神经网络学习波束权重。
- 有方向引导时波束图更平滑稳定,能持续跟踪说话人方向,无引导则模糊漂移。
- 适用于可穿戴设备的双耳语音增强,保留左右耳间的时间和强度差异。
本文提出一种面向动态声学环境的深度波束成形框架,通过深度神经网络从多通道噪声信号中学习时变波束权重,并由移动目标说话人的连续追踪相对传递函数(RTF)引导。我们利用8麦克风线性阵列分析网络在三种模式下的空间行为:(i) 真实RTF作为理想引导,(ii) 子空间追踪的RTF估计值引导,(iii) 无RTF引导。结果表明,采用RTF引导的模型生成更平滑、空间一致性更高的窄带与宽带波束图,能有效跟踪目标到达方向(DOA),而无引导模型无法维持清晰的空间聚焦。进一步将该框架扩展至双耳波束成形,用于动态说话人增强。系统基于头相关传递函数(HRTF)模拟移动声源的声学场景,实现左右耳的逼真空间渲染。通过间耳级差(ILD)和间耳时差(ITD)量化评估空间线索保留效果,验证其在可听设备应用中的适用性。
原文摘要 · Abstract (English)
In this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a continuously tracked relative transfer function (RTF) of a moving target speaker. We analyze the network's spatial behavior on an 8-microphone linear array by evaluating narrowband and wideband beampatterns in three modes: (i) oracle guidance with true RTFs, (ii) guidance with subspace-tracked RTF estimates, and (iii) operation without RTF guidance. Results show that RTF guidance yields smoother, more spatially consistent beampatterns that track the target direction of arrival (DOA), whereas the unguided model fails to maintain a clear spatial focus. We further extend the framework to binaural beamforming for dynamic target-speaker enhancement. The system is trained using a head-related transfer function (HRTF)-based acoustic simulation of a moving source, enabling realistic spatial rendering at the left and right ears. Spatial cue preservation is quantitatively evaluated in terms of interaural level differences (ILD) and interaural time differences (ITD), demonstrating the method's suitability for hearable applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。