arXiv:2606.21227cs.SD2026-06中稿 · InterSpeech 2026, …

用自然语言直接编辑音频,无需训练数据

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption

论文配图:Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption
图 1 · 摘自论文原文
  • 将音频转为语义描述,通过改写描述来生成新音频
  • 零样本下实现自由文本控制的音频编辑,效果媲美专业模型
  • 适合语音、音乐、音效等多场景,无需预设模板

现有文本引导的音频编辑方法依赖成对训练数据、预定义操作模板,并在语音、音乐和音效间使用独立处理流程。本文提出Bagpiper-Edit,实现通过自由形式自然语言指令进行开放式音频编辑。我们将音频编辑重构为丰富描述(rich-caption)重写任务,将音频片段视为语义表示。用户请求被转化为修改后的描述,进而指导模型以原始音频为声学上下文锚点生成目标编辑音频。该方法突破了成对数据依赖,实现强大的零样本编辑能力。在语音、音频及自由编辑任务上的评估显示,Bagpiper-Edit 在保持与原音频高一致性的同时,在多数情况下达到与其他专家模型相当的性能。

原文摘要 · Abstract (English)

Current text-guided audio editing methods rely on paired training data, predefined operation templates, and separate processing pipelines across speech, music, and sound. We present Bagpiper-Edit to enable open-ended audio editing via free-form natural language instructions. We reformulate audio editing as a rich-caption rewriting task by treating a rich caption as the semantic representation of an audio clip. The user request is translated into an edited caption, which then guides Bagpiper-Edit to generate the target edited audio with the original audio as contextual acoustic anchor. This unlocks the potential of free-form editing, and circumvents the need for paired audio-editing training data, enabling powerful zero-shot editing capabilities. Evaluations across speech, audio, and free-form editing show Bagpiper-Edit maintains good consistency to the original audio and achieves similar performance to other expert models in most cases. Demo: https://bagpiper-edit.github.io, Codes: https://github.com/espnet/espnet/pull/6417 & https://github.com/HsunGong/espnet

音频编辑零样本自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。