无需提示词,用反向生成融合声音,创造新听感。
Abstract Sound Fusion with Unconditional Inversion Models
- 用SDE/ODE反演模型逆推采样过程,保留原始声特征。
- 融合后声音效果超越简单叠加,具备全新听觉特质。
- 无需文本提示,适合音效设计与创意音频生成。
抽象声音是指听众无法识别其对应真实世界声事件的声音。声音融合旨在将原声与参考声结合,生成一个具有超越两者简单叠加的听觉特性的新声音。为此,我们采用反演技术,在保持原始样本关键特征的同时实现可控合成。本文提出基于DPMSolver++采样器的新颖SDE和ODE反演模型,通过将模型输出设为常量来逆转采样过程,消除噪声预测项带来的循环依赖。该反演方法无需提示词条件,同时在采样过程中仍可灵活引导。
原文摘要 · Abstract (English)
An abstract sound is defined as a sound that does not disclose identifiable real-world sound events to a listener. Sound fusion aims to synthesize an original sound and a reference sound to generate a novel sound that exhibits auditory features beyond mere additive superposition of the sound constituents. To achieve this fusion, we employ inversion techniques that preserve essential features of the original sample while enabling controllable synthesis. We propose novel SDE and ODE inversion models based on DPMSolver++ samplers that reverse the sampling process by configuring model outputs as constants, eliminating circular dependencies incurred by noise prediction terms. Our inversion approach requires no prompt conditioning while maintaining flexible guidance during sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。