arXiv:2507.19369eess.AScs.SD2025-07被引 1

利用个体化耳部传声函数,从多人混杂声音中精准提取目标说话人。

Binaural Target Speaker Extraction using Individualized HRTF

  • 基于个体化头相关传输函数,直接处理复数短时傅里叶变换信号。
  • 在混响环境下仍保持语音清晰度与方向性,优于现有方法。
  • 适合需要真实空间听感的耳机语音增强场景。

本文针对多说话人同时存在时的双耳目标说话人分离问题,提出一种新方法,利用个体听众的头相关传输函数(HRTF)实现目标说话人的分离。该方法无需依赖说话人嵌入,具有说话人无关性。采用全复数神经网络,直接作用于混合音频信号的复数短时傅里叶变换(STFT)表示,并与实虚部分开处理的网络进行对比,验证了前者的优越性。首先在无混响、无噪声条件下评估,成功实现优异的分离性能并完整保留目标信号的双耳线索。随后扩展至混响环境,方法表现出强鲁棒性,既维持语音清晰度和声源方向性,又有效抑制混响。与现有双耳目标说话人提取方法的对比分析表明,该方法在降噪和感知质量上达到当前最优水平,且在保持双耳线索方面具有显著优势。

原文摘要 · Abstract (English)

In this work, we address the problem of binaural target-speaker extraction in the presence of multiple simultane-ous talkers. We propose a novel approach that leverages the individual listener's Head-Related Transfer Function (HRTF) to isolate the target speaker. The proposed method is speaker-independent, as it does not rely on speaker embeddings. We employ a fully complex-valued neural network that operates directly on the complex-valued Short-Time Fourier transform (STFT) of the mixed audio signals, and compare it to a Real-Imaginary (RI)-based neural network, demonstrating the advantages of the former. We first evaluate the method in an anechoic, noise-free scenario, achieving excellent extraction performance while preserving the binaural cues of the target signal. We then extend the evaluation to reverberant conditions. Our method proves robust, maintaining speech clarity and source directionality while simultaneously reducing reverberation. A comparative analysis with existing binaural Target Speaker Extraction (TSE) methods shows that the proposed approach achieves performance comparable to state-of-the-art techniques in terms of noise reduction and perceptual quality, while providing a clear advantage in preserving binaural cues. Demo-page: https://bi-ctse-hrtf.github.io

语音分离双耳处理听觉感知复数网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。