用神经网络优化耳机3D音效,提升高频频段定位精度。
Ambisonics Binaural Rendering via Masked Magnitude Least Squares
- 用神经网络结合频谱掩码优化低阶声场系数
- 高频频段耳间相位保真度提升,中线定位更准
- 适合麦克风阵列有限或带宽受限的场景
头相关传输函数(HRTFs)的低阶表示在耳机3D音频渲染中仍具挑战。现有幅度最小二乘法(MagLS)忽略高频耳间相位信息以降低幅度误差。本文提出掩码幅度最小二乘法(Masked Magnitude Least Squares),利用神经网络优化Ambisonics系数,并引入时空频谱加权掩码,控制幅度重建精度。实验表明,该方法有效保留了低阶HRTFs中的高频凹陷特征,相比MagLS显著改善了中平面定位性能,且对整体幅度重建精度影响微小。
原文摘要 · Abstract (English)
Ambisonics rendering has become an integral part of 3D audio for headphones. It works well with existing recording hardware, the processing cost is mostly independent of the number of sound sources, and it elegantly allows for rotating the scene and listener. One challenge in Ambisonics headphone rendering is to find a perceptually well behaved low-order representation of the Head-Related Transfer Functions (HRTFs) that are contained in the rendering pipe-line. Low-order rendering is of interest, when working with microphone arrays containing only a few sensors, or for reducing the bandwidth for signal transmission. Magnitude Least Squares rendering became the de facto standard for this, which discards high-frequency interaural phase information in favor of reducing magnitude errors. Building upon this idea, we suggest Masked Magnitude Least Squares, which optimized the Ambisonics coefficients with a neural network and employs a spatio-spectral weighting mask to control the accuracy of the magnitude reconstruction. In the tested case, the weighting mask helped to maintain high-frequency notches in the low-order HRTFs and improved the modeled median plane localization performance in comparison to MagLS, while only marginally affecting the overall accuracy of the magnitude reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。