用神经网络动态生成随头部移动的个人声区,实时高效且隔离效果好。
SANN-PSZ: Spatially Adaptive Neural Network for Head-Tracked Personal Sound Zones
- 输入头位坐标,输出自适应声区滤波系数
- 实测环境隔离效果相当或更好,滤波伪影更少
- 相比传统方法压缩100倍、提速10倍,适合实时应用
提出一种基于深度学习的框架,通过空间自适应神经网络(SANN)实现头追踪下的动态个人声区(PSZ)渲染。SANN 输入听众头部坐标,输出声区滤波系数。模型可使用模拟声学传输函数(ATFs)并结合数据增强训练,以提升在不确定环境中的鲁棒性;或在已知条件下混合使用模拟与实测ATFs进行定制化训练。研究发现,在训练数据中增强房间混响比增强系统缺陷更能有效提升模型鲁棒性;在损失函数中加入滤波器紧凑性约束对性能影响不大。与传统滤波设计方法对比显示,在无实测ATFs时,该模型在真实环境中达到同等或更高隔离度,且滤波伪影更少。此外,模型实现100倍的数据压缩和10倍的计算效率提升,适用于实时适应头部运动的声区渲染。
原文摘要 · Abstract (English)
A deep learning framework for dynamically rendering personal sound zones (PSZs) with head tracking is presented, utilizing a spatially adaptive neural network (SANN) that inputs listeners' head coordinates and outputs PSZ filter coefficients. The SANN model is trained using either simulated acoustic transfer functions (ATFs) with data augmentation for robustness in uncertain environments or a mix of simulated and measured ATFs for customization under known conditions. It is found that augmenting room reflections in the training data can more effectively improve the model robustness than augmenting the system imperfections, and that adding constraints such as filter compactness to the loss function does not significantly affect the model's performance. Comparisons of the best-performing model with traditional filter design methods show that, when no measured ATFs are available, the model yields equal or higher isolation in an actual room environment with fewer filter artifacts. Furthermore, the model achieves significant data compression (100x) and computational efficiency (10x) compared to the traditional methods, making it suitable for real-time rendering of PSZs that adapt to the listeners' head movements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。