arXiv:2509.14003cs.SDcs.AI2025-09中稿 · ICASSP 2026被引 9

用流匹配方法实现精准文本引导的音频编辑,无需额外标注。

RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing

  • 基于修正流匹配构建端到端音频编辑框架
  • 在复杂重叠事件音频上实现语义对齐且无需辅助标签
  • 适合需要高精度音频修改的研究与应用

扩散模型在文本到音频生成方面取得了显著进展,但文本引导的音频编辑仍处于起步阶段。该任务要求在保留音频其余部分的前提下,精确修改目标内容,需根据文本提示进行精准定位和忠实编辑。现有基于训练或零样本的方法依赖完整字幕或高成本优化,难以应对复杂编辑场景且实用性不足。本文提出一种新颖的端到端高效修正流匹配扩散框架用于音频编辑,并构建了一个包含重叠多事件音频的数据集,以支持复杂场景下的训练与评估。实验表明,该模型在无需辅助字幕或掩码的情况下实现忠实的语义对齐,且在各项指标上保持竞争力。

原文摘要 · Abstract (English)

Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the target content within an audio signal while preserving the rest, thus demanding precise localization and faithful editing according to the text prompt. Existing training-based and zero-shot methods that rely on full-caption or costly optimization often struggle with complex editing or lack practicality. In this work, we propose a novel end-to-end efficient rectified flow matching-based diffusion framework for audio editing, and construct a dataset featuring overlapping multi-event audio to support training and benchmarking in complex scenarios. Experiments show that our model achieves faithful semantic alignment without requiring auxiliary captions or masks, while maintaining competitive editing quality across metrics.

音频编辑扩散模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。