用自然语言自由编辑音频,无需预设指令
SAO-Instruct: Free-form Audio Editing using Natural Language Instructions
- 基于Stable Audio Open,支持任意自由文本指令编辑音频
- 在真实音频和未见指令上表现良好,主观评测优于现有方法
- 开源代码与模型,推动语音编辑研究发展
生成模型已在根据简短文本描述合成高质量音频方面取得显著进展。然而,使用自然语言编辑现有音频仍鲜有探索。当前方法要么需要完整描述编辑后音频,要么受限于预定义指令,灵活性不足。本文提出SAO-Instruct,基于Stable Audio Open,可对音频片段进行任意自由形式的自然语言编辑。为训练模型,我们利用Prompt-to-Prompt、DDPM反演及人工编辑流程构建了音频编辑三元组数据集(输入音频、编辑指令、输出音频)。尽管部分在合成数据上训练,模型在真实场景音频和未见指令上仍具有良好泛化能力。实验表明,SAO-Instruct在客观指标上表现竞争力,并在主观听感测试中超越其他音频编辑方法。为促进后续研究,我们开源代码与模型权重。
原文摘要 · Abstract (English)
Generative models have made significant progress in synthesizing high-fidelity audio from short textual descriptions. However, editing existing audio using natural language has remained largely underexplored. Current approaches either require the complete description of the edited audio or are constrained to predefined edit instructions that lack flexibility. In this work, we introduce SAO-Instruct, a model based on Stable Audio Open capable of editing audio clips using any free-form natural language instruction. To train our model, we create a dataset of audio editing triplets (input audio, edit instruction, output audio) using Prompt-to-Prompt, DDPM inversion, and a manual editing pipeline. Although partially trained on synthetic data, our model generalizes well to real in-the-wild audio clips and unseen edit instructions. We demonstrate that SAO-Instruct achieves competitive performance on objective metrics and outperforms other audio editing approaches in a subjective listening study. To encourage future research, we release our code and model weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。