用大模型统一处理多种音频编辑任务,精准又保真。
MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model
- 基于音频语言模型,统一支持添加、替换、删除等五类编辑
- 构建大规模细粒度标注数据集,提升指令跟随与定位精度
- 适合需要高保真音频编辑的AI音乐创作与内容生产场景
文本引导的音频编辑旨在修改特定声学事件,同时严格保留非目标内容。尽管已有进展,现有方法仍存在根本局限:无需训练的方法常因扩散反演导致信号退化,而需训练的方法受限于高质量配对数据稀缺,且任务定义仅覆盖少数编辑操作。此外,标准架构通常将文本与音频处理解耦,难以实现指令与具体声学上下文的对齐。为此,我们提出MMEdit,一种由音频语言模型驱动的统一音频编辑框架。系统性扩展任务定义,涵盖添加、替换、移除、重排及属性修改等五类操作;设计可扩展的数据合成管道,构建大规模带细粒度事件级标注的配对数据集;集成Qwen2-Audio编码器与基于MMDiT的生成器,实现精确跨模态对齐与局部编辑。实验表明,该方法在编辑定位精度、指令遵循能力及未编辑区域保真度方面均表现优异。
原文摘要 · Abstract (English)
Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally limited. Training-free methods often suffer from signal degradation caused by diffusion inversion, while training-based methods, although achieving higher generation quality, are severely constrained by the scarcity of high-quality paired data and task formulations that cover only a narrow subset of editing operations. In addition, standard architectures typically decouple text and audio processing, limiting the ability to align instructions with specific acoustic contexts. To address these challenges, we propose MMEdit, an audio-language-model-driven framework for unified audio editing. We systematically extend task definitions to cover a comprehensive range of editing operations, including addition, replacement, removal, reordering, and attribute modification. Furthermore, we design a scalable data synthesis pipeline to construct large-scale paired datasets with fine-grained event-level annotations. To capture complex editing semantics, we integrate a Qwen2-Audio encoder with an MMDiT-based generator, enabling precise cross-modal alignment and localized editing. Experimental results demonstrate that our method achieves superior editing localization accuracy, robust instruction following, and high fidelity in non-edited regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。