用物理规律指导生成,让语音位置更真实可控。
PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation

- 共享路径点描述统一语言与轨迹控制
- 引入方向与距离的物理一致性约束
- 适合游戏影视音效设计人员使用
文本到空间音频生成(如文本到一阶全向声学,FOA)为数十亿美元的游戏与影视产业提供了便捷的音频创作方式。然而,现有方法多依赖数据驱动,可能生成违反声源方向与距离间声学关系的音频;且将描述性控制与参数化控制分离,导致用户需在易用性与精度间权衡。本文提出 PhysWave,一种基于物理引导的潜在扩散模型,用于可控文本到FOA音频生成。PhysWave通过共享的路径点-描述表示统一自然语言与轨迹控制,并在扩散训练中引入两个可微分的声学先验:球谐方向一致性与反平方距离一致性。为支持动态空间生成,我们构建了一个包含300K片段、涵盖多种声音类别和声源轨迹的大型FOA数据集。大量实验表明,所提先验能有效提升生成音频的空间一致性,同时保持优异的音频质量。进一步分析显示,这些物理先验在训练阶段增强空间一致性,亦可在推理阶段无需重新训练实现空间精修。
原文摘要 · Abstract (English)
Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acoustic relations between source direction and distance. They also separate descriptive and parametric control, forcing users to trade usability for precision. In this paper, we present PhysWave, a physics-guided latent diffusion model for controllable text-to-FOA generation. PhysWave unifies natural-language and trajectory control through a shared waypoint-caption representation, and augments diffusion training with two differentiable acoustic priors: spherical-harmonic direction consistency and inverse-square distance consistency. To support dynamic spatial generation, we further construct a 300K-clip FOA dataset with diverse sound categories and source trajectories. Extensive results show that the proposed priors help PhysWave generate spatially consistent FOA audio while maintaining competitive audio quality. Further analyses show that these physics priors improve spatial consistency during training and can also be used as inference-time guidance for training-free spatial refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。