让声音在复杂空间中动态回响更真实,且适合可穿戴设备运行。
Differentiable Grouped Feedback Delay Networks for Learning Coupled Volume Acoustics
- 用可微分的分组反馈延时网络,自动匹配多斜率衰减的声学环境。
- 训练后能插值未测量位置,计算量比传统方法低十倍以上。
- 适合移动源和听者在扩展现实中的实时沉浸式音效渲染。
在复杂声学空间中为移动声源和听者渲染动态混响对提升扩展现实(XR)应用的沉浸感至关重要。捕捉空间变化的房间冲激响应(RIRs)成本高且不切实际。而使用实测RIRs进行动态卷积计算开销大,内存需求高,通常无法在可穿戴设备上运行。分组反馈延时网络(GFDNs)可高效渲染耦合声学效果,但其参数需调优以匹配特定空间的混响特征。本文提出可微分的GFDNs(DiffGFDNs),通过优化一组实测多斜率衰减空间的晚混响特征,获得可调参数。训练完成后,DiffGFDN可在未测量位置实现插值。采用并行处理管道,每个八度频段独立处理,参数在推理时可快速更新。在三个耦合房间的RIR数据集上评估,与通用斜率(CS)模型相比,该架构在低内存和计算量下生成多斜率晚混响,能量衰减曲线误差(EDR)更低,八度带能量衰减曲线(EDC)误差略高,且每样本浮点运算量减少一个数量级。
原文摘要 · Abstract (English)
Rendering dynamic reverberation in a complicated acoustic space for moving sources and listeners is challenging but crucial for enhancing user immersion in extended-reality (XR) applications. Capturing spatially varying room impulse responses (RIRs) is costly and often impractical. Moreover, dynamic convolution with measured RIRs is computationally expensive with high memory demands, typically not available on wearable computing devices. Grouped Feedback Delay Networks (GFDNs), on the other hand, allow efficient rendering of coupled room acoustics. However, its parameters need to be tuned to match the reverberation profile of a coupled space. In this work, we propose the concept of Differentiable GFDNs (DiffGFDNs), which have tunable parameters that are optimised to match the late reverberation profile of a set of RIRs captured from a space that exhibits multi-slope decay. Once trained on a finite set of measurements, the DiffGFDN interpolates to unmeasured locations in the space. We propose a parallel processing pipeline that has multiple DiffGFDNs with frequency-independent parameters processing each octave band. The parameters of the DiffGFDN can be updated rapidly during inferencing as sources and listeners move. We evaluate the proposed architecture against the Common Slopes (CS) model on a dataset of RIRs for three coupled rooms. The proposed architecture generates multi-slope late reverberation with low memory and computational requirements, achieving a better energy decay relief (EDR) error and slightly worse octave-band energy decay curve (EDC) errors compared to the CS model. Furthermore, DiffGFDN requires an order of magnitude fewer floating-point operations per sample than the CS renderer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。