用生成模型实现音频的前后扩展与无缝融合,提升音效设计效率。
Generative Audio Extension and Morphing
- 通过掩码噪声潜变量并使用新型无分类器引导,实现音频扩展与变形。
- 生成音频在FAD指标上接近真实样本,主观评测获积极反馈。
- 可微调适配不同静音音频数据,减少幻觉,适合音效设计师使用。
在音频创作任务中,音效设计师常需扩展和融合音频库中的声音。生成式音频模型可通过参考样本生成音频,提供有效解决方案。通过掩码DiT模型的噪声潜变量,并在该掩码潜变量上应用一种新型无分类器引导方法,我们证明:(i) 给定一个音频参考,可按指定时长向前或向后扩展;(ii) 给定两个音频参考,可无缝融合生成目标时长的过渡音频。此外,通过在不同类型的静音音频数据上微调模型,可缓解潜在的幻觉问题。该方法在客观指标上表现优异,生成音频的弗雷谢音频距离(FAD)与训练数据的真实样本相当。主观听觉测试也显示受试者对生成结果评价积极。该技术为更可控、更具表现力的生成式音频框架铺平道路,使音效设计师能摆脱重复性劳动,专注创意过程。
原文摘要 · Abstract (English)
In audio-related creative tasks, sound designers often seek to extend and morph different sounds from their libraries. Generative audio models, capable of creating audio using examples as references, offer promising solutions. By masking the noisy latents of a DiT and applying a novel variant of classifier-free guidance on such masked latents, we demonstrate that: (i) given an audio reference, we can extend it both forward and backward for a specified duration, and (ii) given two audio references, we can morph them seamlessly for the desired duration. Furthermore, we show that by fine-tuning the model on different types of stationary audio data we mitigate potential hallucinations. The effectiveness of our method is supported by objective metrics, with the generated audio achieving Fréchet Audio Distances (FADs) comparable to those of real samples from the training data. Additionally, we validate our results through a subjective listener test, where subjects gave positive ratings to the proposed model generations. This technique paves the way for more controllable and expressive generative sound frameworks, enabling sound designers to focus less on tedious, repetitive tasks and more on their actual creative process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。