根据文字描述生成随时间移动的立体声音效,实现动态听觉体验。
Text2Move: Text-to-moving sound generation via trajectory prediction and temporal alignment
- 用文本预测声音源在三维空间中的运动轨迹
- 训练模型使音频与轨迹精准对齐,生成真实感立体声
- 可嵌入现有语音生成流程,适合影视音效创作
人类听觉感知依赖于三维空间中移动的声音源,但以往生成式声音建模大多局限于单声道或静态空间音频。本文提出一种可控的文本到移动声音生成框架。为支持训练,构建了一个合成数据集,包含双耳格式的移动声音、其空间轨迹以及描述声音事件和空间运动的文字标注。基于该数据集,训练一个文本到轨迹预测模型,可从文本提示中输出声音源的三维运动轨迹。为生成空间音频,先微调预训练的文本到音频生成模型,使其输出与轨迹时间对齐的单声道音频,再利用预测的轨迹模拟空间音频。实验表明,该模型具备合理的空间理解能力。该方法可轻松集成至现有文本到音频生成工作流,并可扩展至其他空间音频格式。
原文摘要 · Abstract (English)
Human auditory perception is shaped by moving sound sources in 3D space, yet prior work in generative sound modelling has largely been restricted to mono signals or static spatial audio. In this work, we introduce a framework for generating moving sounds given text prompts in a controllable fashion. To enable training, we construct a synthetic dataset that records moving sounds in binaural format, their spatial trajectories, and text captions about the sound event and spatial motion. Using this dataset, we train a text-to-trajectory prediction model that outputs the three-dimensional trajectory of a moving sound source given text prompts. To generate spatial audio, we first fine-tune a pre-trained text-to-audio generative model to output temporally aligned mono sound with the trajectory. The spatial audio is then simulated using the predicted temporally-aligned trajectory. Experimental evaluation demonstrates reasonable spatial understanding of the text-to-trajectory model. This approach could be easily integrated into existing text-to-audio generative workflow and extended to moving sound generation in other spatial audio formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。