首个可生成动态3D声音景观的扩散模型,支持语言与轨迹控制。
Generating Moving 3D Soundscapes with Latent Diffusion Models
- 基于潜在扩散框架,实现从文本或轨迹参数生成带运动的3D音频。
- 在百万级模拟数据上训练,空间定位误差低,音质媲美顶尖文本转音频系统。
- 适合虚拟现实、游戏等需要精准移动声源的应用场景。
空间音频在虚拟现实、增强现实、电影和音乐等沉浸式应用中日益重要。现有生成音频模型多局限于单声道或立体声,无法捕捉一阶全向声学(FOA)中的完整三维定位信息。近期的FOA模型虽扩展至文本到音频生成,但仍限于静态声源。本文提出SonicMotion,首个端到端的潜在扩散框架,可生成带有显式运动控制的FOA音频。该框架有两种变体:1)基于自然语言提示的描述性模型;2)结合文本与空间轨迹参数的参数化模型,实现更高精度。为支持训练与评估,构建了一个包含超一百万对模拟FOA音频-描述数据集,涵盖静态与动态声源,并标注方位角、仰角及运动属性。实验表明,SonicMotion在语义对齐方面达到当前最佳水平,感知质量接近领先文本到音频系统,且独特地实现了低空间定位误差。
原文摘要 · Abstract (English)
Spatial audio has become central to immersive applications such as VR/AR, cinema, and music. Existing generative audio models are largely limited to mono or stereo formats and cannot capture the full 3D localization cues available in first-order Ambisonics (FOA). Recent FOA models extend text-to-audio generation but remain restricted to static sources. In this work, we introduce SonicMotion, the first end-to-end latent diffusion framework capable of generating FOA audio with explicit control over moving sound sources. SonicMotion is implemented in two variations: 1) a descriptive model conditioned on natural language prompts, and 2) a parametric model conditioned on both text and spatial trajectory parameters for higher precision. To support training and evaluation, we construct a new dataset of over one million simulated FOA caption pairs that include both static and dynamic sources with annotated azimuth, elevation, and motion attributes. Experiments show that SonicMotion achieves state-of-the-art semantic alignment and perceptual quality comparable to leading text-to-audio systems, while uniquely attaining low spatial localization error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。