arXiv:2509.17219cs.SDcs.LG2025-09

无需反演,用新采样法实现快速高质量文本音频编辑

Virtual Consistency for Audio Editing

  • 通过调整扩散模型采样过程实现无反演编辑
  • 速度比现有方法快,且质量不下降(16人用户测试验证)
  • 适用于任意模型,无需微调或改架构

尽管基于反演的神经方法取得进展,自由形式的文本音频编辑仍是持续挑战。当前方法依赖缓慢的反演过程,限制了实用性。我们提出一种基于虚拟一致性(Virtual Consistency)的音频编辑系统,通过调整扩散模型的采样过程绕过反演。该流程模型无关,无需微调或结构改动,相比近期神经编辑基线实现显著提速。关键的是,该方法在不牺牲质量的前提下实现高效,经量化基准测试和16名参与者的用户研究验证。

原文摘要 · Abstract (English)

Free-form, text-based audio editing remains a persistent challenge, despite progress in inversion-based neural methods. Current approaches rely on slow inversion procedures, limiting their practicality. We present a virtual-consistency based audio editing system that bypasses inversion by adapting the sampling process of diffusion models. Our pipeline is model-agnostic, requiring no fine-tuning or architectural changes, and achieves substantial speed-ups over recent neural editing baselines. Crucially, it achieves this efficiency without compromising quality, as demonstrated by quantitative benchmarks and a user study involving 16 participants.

音频编辑扩散模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。