用深度神经网络实时识别旋转音箱的听觉响应,提升定位精度。
DNN based HRIRs Identification with a Continuously Rotating Speaker Array
- 通过序列到序列学习建模声源旋转时的耳部响应变化。
- 在45°/s转速下,对准误差降低7dB,频谱失真低于2dB。
- 适合需要高精度动态声场建模的虚拟现实与音频系统开发。
传统静态测量头相关脉冲响应(HRIRs)需逐角重置音箱阵列,耗时较长。现有动态方法虽采用连续旋转音箱,但在高速旋转下精度显著下降。为此,本文提出基于深度神经网络的序列建模方法,利用全连接网络捕捉HRIR过渡特性,并引入更新门与重置门实现整段序列的响应识别。模型基于瞬时平方误差(ISE)梯度更新HRIR系数,同时设计可学习的归一化机制,稳定ISE梯度随时间的变化。此外,提出整体序列优化训练策略以防止过拟合。仿真使用FABIAN数据库验证,相较以往解析模型,NM提升超7dB,LSD保持低于2dB(45°/s转速)。实验采用自建音箱阵列,结果表明该方法能有效保留与静态测量一致的精确定位线索。源代码见https://github.com/byko0810/DNN-based-HRIRs-identification。
原文摘要 · Abstract (English)
Conventional static measurement of head-related impulse responses (HRIRs) is time-consuming due to the need for repositioning a speaker array for each azimuth angle. Dynamic approaches using analytical models with a continuously rotating speaker array have been proposed, but their accuracy is significantly reduced at high rotational speeds. To address this limitation, we propose a DNN-based HRIRs identification using sequence-to-sequence learning. The proposed DNN model incorporates fully connected (FC) networks to effectively capture HRIR transitions and includes reset and update gates to identify HRIRs over a whole sequence. The model updates the HRIRs vector coefficients based on the gradient of the instantaneous square error (ISE). Additionally, we introduce a learnable normalization process based on the speaker excitation signals to stabilize the gradient scale of ISE across time. A training scheme, referred to as whole-sequence updating and optimization scheme, is also introduced to prevent overfitting. We evaluated the proposed method through simulations and experiments. Simulation results using the FABIAN database show that the proposed method outperforms previous analytic models, achieving over 7 dB improvement in normalized misalignment (NM) and maintaining log spectral distortion (LSD) below 2 dB at a rotational speed of 45°/s. Experimental results with a custom-built speaker array confirm that the proposed method successfully preserved accurate sound localization cues, consistent with those from static measurement. Source code is available at https://github.com/byko0810/DNN-based-HRIRs-identification
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。