从声音推断房间声学特征,生成任意位置的沉浸式音效
Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
- 用对比学习将声音编码为仅含房间信息的低维特征
- 基于扩散模型生成新位置的房间冲击响应
- 支持任意声源-接收位置组合,适合AR音效渲染
在增强现实(AR)音频中,了解用户真实的声学环境对于使虚拟声音自然融合至关重要。由于实际应用中通常无法进行声学测量,需从现有声源推断房间信息,进而使新增声源具备相同的房间声学特性。关键挑战在于这些新增声源的位置与用于估计的声源不同。本文提出一种基于对比损失训练的编码器网络,将输入声音映射到仅包含房间特性的低维特征空间。随后,采用基于扩散模型的空间房间冲击响应生成器,根据该特征空间和新的声源-接收器位置生成对应响应。结果表明,最终输出同时考虑了房间特性和位置特异性参数。
原文摘要 · Abstract (English)
For audio in augmented reality (AR), knowledge of the users' real acoustic environment is crucial for rendering virtual sounds that seamlessly blend into the environment. As acoustic measurements are usually not feasible in practical AR applications, information about the room needs to be inferred from available sound sources. Then, additional sound sources can be rendered with the same room acoustic qualities. Crucially, these are placed at different positions than the sources available for estimation. Here, we propose to use an encoder network trained using a contrastive loss that maps input sounds to a low-dimensional feature space representing only room-specific information. Then, a diffusion-based spatial room impulse response generator is trained to take the latent space and generate a new response, given a new source-receiver position. We show how both room- and position-specific parameters are considered in the final output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。