用空间参数生成沉浸式3D音效,让声音位置与环境更真实。
ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model
- 基于空间、时间、环境参数生成第一阶全向声场音频。
- 生成音效在空间定位上符合用户设定条件,质量优秀。
- 适合影视游戏音效生成,也适用于虚拟现实场景构建。
我们提出 ImmerseDiffusion,一个端到端的生成式音频模型,可依据声源的空间、时间及环境条件生成三维沉浸式音景。该模型训练生成第一阶全向声场(FOA)音频,一种包含四个通道的常规空间音频格式,可渲染为多声道输出。系统由空间音频编解码器组成,将FOA音频映射至潜在空间;同时使用多种用户输入类型(如文本提示、空间参数、时间参数、环境声学参数)训练潜在扩散模型,并可选加入基于对比语言与音频预训练(CLAP)风格的时空音频和文本编码器。我们设计了评估指标以衡量生成音频的质量与空间一致性。实验对比了两种模式:‘描述型’(使用空间文本提示)与‘参数型’(使用非空间文本提示加空间参数)。评估结果表明,生成内容与用户设定条件高度一致,具备可靠的声场保真度。
原文摘要 · Abstract (English)
We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sound objects. ImmerseDiffusion is trained to generate first-order ambisonics (FOA) audio, which is a conventional spatial audio format comprising four channels that can be rendered to multichannel spatial output. The proposed generative system is composed of a spatial audio codec that maps FOA audio to latent components, a latent diffusion model trained based on various user input types, namely, text prompts, spatial, temporal and environmental acoustic parameters, and optionally a spatial audio and text encoder trained in a Contrastive Language and Audio Pretraining (CLAP) style. We propose metrics to evaluate the quality and spatial adherence of the generated spatial audio. Finally, we assess the model performance in terms of generation quality and spatial conformance, comparing the two proposed modes: ``descriptive", which uses spatial text prompts) and ``parametric", which uses non-spatial text prompts and spatial parameters. Our evaluations demonstrate promising results that are consistent with the user conditions and reflect reliable spatial fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。