用AI加速声音传播模拟,提升语音处理效果
Efficient learning-based sound propagation for virtual and real-world audio processing applications
- 基于学习的RIR生成器比传统方法快100倍,支持多种场景和音频格式
- 从语音和视觉信号中估计混响响应,使远场语音识别错误率降低6.9%
- 结合真实数据生成新声学环境,适合虚拟音效、配音等应用
声音传播描述声波在介质(如空气)中传播的过程,其特性由房间脉冲响应(RIR)表征,受声源与听者位置、房间几何形状及材料影响。传统物理模拟器虽能准确计算特定环境下的RIR,但存在效率瓶颈。为此,本文提出三项创新:首先,设计一种学习型RIR生成器,速度较交互式射线追踪模拟器提升两个数量级,可直接输入统计与传统参数,生成单声道与双声道RIR,适用于重建与合成的3D场景,且在语音识别(ASR)、语音增强与分离任务中表现优于物理模拟器;其次,提出无需3D环境表示即可从混响语音与视觉线索中估计RIR,用于训练数据增强,使远场ASR词错误率下降6.9%;该方法还支持视觉声学匹配、新视角声学合成与语音配音,并通过感知评估验证;最后,引入IR-GAN,利用真实RIR学习声学参数,生成模拟不同声学环境的新RIR,其在远场ASR基准上比射线追踪模拟器提升8.95%。
原文摘要 · Abstract (English)
Sound propagation is the process by which sound energy travels through a medium, such as air, to the surrounding environment as sound waves. The room impulse response (RIR) describes this process and is influenced by the positions of the source and listener, the room's geometry, and its materials. Physics-based acoustic simulators have been used for decades to compute accurate RIRs for specific acoustic environments. However, we have encountered limitations with existing acoustic simulators. To address these limitations, we propose three novel solutions. First, we introduce a learning-based RIR generator that is two orders of magnitude faster than an interactive ray-tracing simulator. Our approach can be trained to input both statistical and traditional parameters directly, and it can generate both monaural and binaural RIRs for both reconstructed and synthetic 3D scenes. Our generated RIRs outperform interactive ray-tracing simulators in speech-processing applications, including ASR, Speech Enhancement, and Speech Separation. Secondly, we propose estimating RIRs from reverberant speech signals and visual cues without a 3D representation of the environment. By estimating RIRs from reverberant speech, we can augment training data to match test data, improving the word error rate of the ASR system. Our estimated RIRs achieve a 6.9% improvement over previous learning-based RIR estimators in far-field ASR tasks. We demonstrate that our audio-visual RIR estimator aids tasks like visual acoustic matching, novel-view acoustic synthesis, and voice dubbing, validated through perceptual evaluation. Finally, we introduce IR-GAN to augment accurate RIRs using real RIRs. IR-GAN parametrically controls acoustic parameters learned from real RIRs to generate new RIRs that imitate different acoustic environments, outperforming Ray-tracing simulators on the far-field ASR benchmark by 8.95%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。