arXiv:2601.06621eess.AScs.SD2026-01被引 2

让多人在不同位置听立体声,互不干扰。

Stereo Audio Rendering for Personal Sound Zones Using a Binaural Spatially Adaptive Neural Network (BSANN)

  • 用神经网络为每只耳朵定制扬声器滤波器
  • 实测三维空间感知提升,串扰降低10.55分贝以上
  • 适合车载、家庭影院等真实声学环境

本文提出一种双耳化个人声区(PSZ)渲染框架,使多名头戴追踪的听众能独立接收立体声音频。现有系统多采用单声道渲染,无法分别控制左右耳,限制了空间成像质量。所提方法利用双耳空间自适应神经网络(BSANN)生成针对每只耳朵优化的扬声器滤波器,在多个听众耳部重建期望声场。框架融合了无混响测量的扬声器频率响应、解析建模的扬声器指向性以及刚性球头相关传递函数(HRTFs),提升了声学精度与空间还原度。此外,显式主动串扰消除(XTC)阶段进一步增强了三维空间感知。实验表明,在100–20,000 Hz频段内,对位隔离度(IZI)、节目隔离度(IPI)和串扰消除(XTC)的对数频率加权值分别达到10.23/10.03 dB、11.11/9.16 dB、10.55/11.13 dB,显著优于现有方法。耳级控制、精确声学建模与集成主动XTC的结合,实现了高隔离性、对房间不对称更强鲁棒性及更真实的声场再现。

原文摘要 · Abstract (English)

A binaural rendering framework for personal sound zones (PSZs) is proposed to enable multiple head-tracked listeners to receive fully independent stereo audio programs. Current PSZ systems typically rely on monophonic rendering and therefore cannot control the left and right ears separately, which limits the quality and accuracy of spatial imaging. The proposed method employs a Binaural Spatially Adaptive Neural Network (BSANN) to generate ear-optimized loudspeaker filters that reconstruct the desired acoustic field at each ear of multiple listeners. The framework integrates anechoically measured loudspeaker frequency responses, analytically modeled transducer directivity, and rigid-sphere head-related transfer functions (HRTFs) to enhance acoustic accuracy and spatial rendering fidelity. An explicit active crosstalk cancellation (XTC) stage further improves three-dimensional spatial perception. Experiments show significant gains in measured objective performance metrics, including inter-zone isolation (IZI), inter-program isolation (IPI), and crosstalk cancellation (XTC), with log-frequency-weighted values of 10.23/10.03 dB (IZI), 11.11/9.16 dB (IPI), and 10.55/11.13 dB (XTC), respectively, over 100-20,000 Hz. The combined use of ear-wise control, accurate acoustic modeling, and integrated active XTC produces a unified rendering method that delivers greater isolation performance, increased robustness to room asymmetry, and more faithful spatial reproduction in real acoustic environments.

声场渲染双耳处理空间音频串扰消除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。