用学习到的声学先验匹配混响语音,实现高效真实感混响渲染。
Matching Reverberant Speech Through Learned Acoustic Embeddings and Feedback Delay Networks
- 通过声学先验将混响参数估计转为信号匹配任务。
- 新提出的反馈延时网络可还原目标空间的频响衰减与直达比。
- 在感知真实度上优于现有自动调参方法,适合音频增强现实应用。
混响传递了环境的重要声学线索,有助于空间感知和沉浸体验。对于听觉增强现实(AAR)系统而言,在缺乏显式声学测量的情况下实现实时生成合理混响仍是一大挑战。本文将人工混响参数的盲估计问题建模为混响信号匹配任务,利用学习到的房间声学先验。此外,提出一种反馈延时网络(FDN)结构,能够再现目标空间的频率相关衰减时间与直达-混响比。实验对比当前领先的自动FDN调参方法,结果表明在估计的声学参数及人工混响语音的感知真实度方面均有提升。该方法展现了在AAR应用中实现高效、感知一致混响渲染的巨大潜力。
原文摘要 · Abstract (English)
Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible reverberation in real time remains a key challenge, especially when explicit acoustic measurements are unavailable. We address this by formulating blind estimation of artificial reverberation parameters as a reverberant signal matching task, leveraging a learned room-acoustic prior. Furthermore, we propose a feedback delay network (FDN) structure that reproduces both frequency-dependent decay times and the direct-to-reverberation ratio of a target space. Experimental evaluation against a leading automatic FDN tuning method demonstrates improvements in estimated room-acoustic parameters and perceptual plausibility of artificial reverberant speech. These results highlight the potential of our approach for efficient, perceptually consistent reverberation rendering in AAR applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。