arXiv:2603.16668eess.AScs.SD2026-03

用头相关传输函数提升双耳语音分离的方位保真度

HRTF-guided Binaural Target Speaker Extraction with Real-World Validation

  • 以听众的头相关传输函数作为空间先验指导语音分离
  • 在真实头体模拟器上验证,显著提升语音清晰度与方位感知
  • 适用于需要精准空间定位的听觉增强场景

本文提出一种基于头相关传输函数(HRTF)引导的双耳目标说话人提取框架,用于从多个并发声源的混合信号中分离目标语音。不同于依赖到达方向或录音信号的传统方法,该方法利用听者的HRTF作为显式空间先验,避免了空间感知失真。框架基于多通道深度盲源分离模型,适配双耳分离任务,并在多样人群测量的HRTF数据上训练,实现跨个体泛化,无需个性化调参。通过结合HRTF衍生的空间信息,该方法在保留双耳线索的同时提升了语音质量与可懂度。性能在仿真与真实头体模拟器(HATS)录制数据上得到验证。

原文摘要 · Abstract (English)

This paper presents a Head-Related Transfer Function (HRTF)-guided framework for binaural Target Speaker Extraction (TSE) from mixtures of concurrent sources. Unlike conventional TSE methods based on Direction of Arrival (DOA) estimation or enrollment signals, which often distort perceived spatial location, the proposed approach leverages the listener's HRTF as an explicit spatial prior. The proposed framework is built upon a multi-channel deep blind source separation backbone, adapted to the binaural TSE setting. It is trained on measured HRTFs from a diverse population, enabling cross-listener generalization rather than subject-specific tuning. By conditioning the extraction on HRTF-derived spatial information, the method preserves binaural cues while enhancing speech quality and intelligibility. The performance of the proposed framework is validated through simulations and real recordings obtained from a head and torso simulator (HATS).

语音分离双耳处理空间感知HRTF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。