arXiv:2409.12466cs.SDeess.AS2024-09被引 23

无需训练即可精准编辑音频,保留原内容不变。

AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework

论文配图:AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
图 1 · 摘自论文原文
  • 基于预训练扩散模型,结合空文本反演与终止信号抑制
  • 实现高保真编辑,未修改部分几乎无失真
  • 适合需要快速音效修改的创作者与研究人员

基于扩散模型的文本到音频(TTA)生成已取得显著进展,利用潜空间扩散模型(LDM)生成高质量、多样化且符合指令的音频。然而,除生成外,音频编辑任务同样重要但关注较少。音频编辑面临两大挑战:精确修改与保留未编辑部分。尽管基于LDM的工作在图像处理中已有效解决这些问题,但在音频领域应用极少。本文提出AudioEditor,一个基于预训练扩散型TTA模型的免训练音频编辑框架。该框架引入空文本反演(Null-text Inversion)与结束时刻抑制(EOT-suppression)方法,使模型在执行精确编辑的同时保持原始音频特征。大量客观与主观实验验证了AudioEditor在生成高质量音频编辑结果方面的有效性。代码与演示可访问 https://github.com/NKU-HLT/AudioEditor。

原文摘要 · Abstract (English)

Diffusion-based text-to-audio (TTA) generation has made substantial progress, leveraging latent diffusion model (LDM) to produce high-quality, diverse and instruction-relevant audios. However, beyond generation, the task of audio editing remains equally important but has received comparatively little attention. Audio editing tasks face two primary challenges: executing precise edits and preserving the unedited sections. While workflows based on LDMs have effectively addressed these challenges in the field of image processing, similar approaches have been scarcely applied to audio editing. In this paper, we introduce AudioEditor, a training-free audio editing framework built on the pretrained diffusion-based TTA model. AudioEditor incorporates Null-text Inversion and EOT-suppression methods, enabling the model to preserve original audio features while executing accurate edits. Comprehensive objective and subjective experiments validate the effectiveness of AudioEditor in delivering high-quality audio edits. Code and demo can be found at https://github.com/NKU-HLT/AudioEditor.

音频编辑扩散模型免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。