无需反演的音频编辑新方法,通过最优传输稳定编辑轨迹。
EchoEdit: Stabilizing Inversion-Free Audio Editing via Optimal Transport Geometry

- 直接差分源与目标提示的漂移构造编辑场,免去反演和优化。
- 在音效与音乐编辑任务中,目标对齐度提升18%,源结构保留更强。
- 适合需要快速、稳定音频编辑的研究者与创作者使用。
基于预训练生成模型的文本引导音频编辑通常依赖反演或加噪。这种架构带来结构性权衡:更强的编辑需更深度破坏应保持不变的节奏、瞬态、音色和长程结构。本文提出EchoEdit,一种无需训练、无需反演的实时音频编辑框架,通过差分源与目标提示条件下的漂移来直接构建编辑场。该方法避免了显式反演、成对编辑数据和测试时优化,但其随机源边缘引入不确定性漂移,导致小扰动沿编辑路径累积,使编辑潜在变量偏离音频数据流形。为此,我们进一步提出EchoEdit+,一种基于最优传输正则化的扩展,通过最小化编辑变量与源条件音频流形间的运输成本来稳定直接编辑路径。所得最优传输耦合压缩了噪声状态下的随机位移,使模型查询更贴近训练分布,同时保留结构信息并支持语义变化。在音效与音乐编辑实验中,EchoEdit+相比基于反演的基线和无正则化直接编辑器,在目标提示对齐度和源保留方面均有显著提升。代码与数据集将公开发布。
原文摘要 · Abstract (English)
Text-guided audio editing with pretrained generative models is commonly implemented through inversion or noising. This topology induces a structural trade-off, as stronger edits require deeper corruption of the very rhythm, transients, timbre, and long-range form that should remain unchanged. Here, we introduce EchoEdit, a training-free and inversion-free framework for real-audio editing that directly constructs an editing field by differencing the drifts conditioned on the source and target prompts. This construction avoids explicit source inversion, paired edit data, and test-time optimization, but its stochastic source marginals introduce uncertainty drift, where small random deviations accumulate along the editing trajectory and can move the edited latent away from the audio data manifold. To address this limitation, we further propose EchoEdit+, an optimal-transport-regularized extension that stabilizes the direct editing path by minimizing the transportation cost between edited variables and the source-conditioned audio manifold. The resulting OT coupling contracts the stochastic displacement at noisy states, keeps model queries closer to the training distribution, and preserves structural information while allowing semantic change. Experiments on sound-effect and music editing demonstrate that EchoEdit+ improves target-prompt alignment and source preservation over inversion-based baselines and the unregularized direct editor. Code and dataset will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。