用分阶段扩散变换器实现高效精准的文本引导音频编辑
RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

- 采用粗到细的两阶段扩散变换器,先全局对齐再局部精修
- 在重叠音事件和复杂指令下显著提升编辑准确率与效率
- 适合需要高精度音频修改的研究者与音频工程师
音频编辑旨在根据文本指令修改现有音频片段中的特定内容,同时保留其余声学信息。尽管扩散模型取得显著进展,现有基于训练的编辑方法主要依赖卷积U-Net主干中的局部归纳偏置和交叉注意力交互,常阻碍长距离语义对齐及指令的精确理解与定位。相比之下,扩散变换器具备更强的全局建模与多模态融合能力,但现有编辑架构通常仅简单堆叠扩散变换器块,所有块中对拼接的音频与文本标记进行联合注意力计算,导致复杂度随标记长度呈二次增长。为平衡编辑性能与效率,我们提出一种基于修正流匹配(Rectified Flow Matching, RFM)的新颖指令引导音频编辑框架——RFM-Editing 2,其基于混合两阶段扩散变换器。该模型在低分辨率阶段对音频与文本标记执行联合注意力以建立粗粒度语义对齐,随后切换至交替使用联合注意力与交叉注意力块,在高分辨率阶段细化编辑细节。这一粗到细策略实现了高效且精准的指令引导音频编辑。实验表明,该框架在涉及重叠音频事件和复杂指令的挑战性任务中取得显著性能提升,同时大幅提高编辑效率。
原文摘要 · Abstract (English)
Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models, existing training-based editing methods mainly rely on the local inductive biases and cross-attention interaction in convolutional U-Net backbones, which often hinder long-range semantic alignment and precise understanding and localization of instructions. In contrast, diffusion transformers provide stronger global modeling and multimodal fusion, but existing editing architectures usually adopt a simple stack of diffusion transformer blocks. Applying joint attention over concatenated audio and text tokens in all blocks results in quadratic complexity with respect to token length. To balance editing performance and efficiency, we propose a novel instruction-guided audio editing framework based on rectified flow matching (RFM), named RFM-Editing 2, built on a hybrid two-stage diffusion transformer. The proposed model performs joint attention over audio and text tokens to establish coarse semantic alignment at the low-resolution stage, then switches to alternating joint-attention and cross-attention blocks to refine editing details at the high-resolution stage. This coarse-to-fine strategy enables efficient and accurate instruction-guided audio editing. Experiments show that the proposed framework achieves notable performance gains on challenging editing tasks involving overlapping audio events and complex instructions, while substantially improving editing efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。