arXiv:2606.15186cs.SDcs.AI2026-06中稿 · Interspeech 2026被引 1

无需训练即可精准编辑音频,保持背景一致性。

FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing

论文配图:FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing
图 1 · 摘自论文原文
  • 基于TangoFlux模型,通过逆向优化提取目标片段
  • 调度注意力解耦仅修改目标区域,保留原声上下文
  • 支持音频移除与非刚性替换,适合音视频编辑场景

文本到音频(TTA)生成已取得显著进展,但实现精确且一致的音频编辑仍是重大挑战。现有方法难以在时间一致性与背景保留之间取得平衡。本文提出FreeSonic,一个无需训练的框架,利用最先进的基于修正流的TangoFlux模型。FreeSonic采用优化的反演-逆过程和联合文本-音频注意力图,实现目标片段的精确提取。内容编辑方面,提出一种新的调度注意力解耦机制,将修改限制在目标区域,同时保留原始声学上下文。此外,面向任务的噪声注入增强了对音频移除和非刚性替换等任务的适应性。大量实验结果表明,FreeSonic在高保真度和高效性之间实现了优越平衡,为精确且一致的音频编辑提供了优质解决方案。

原文摘要 · Abstract (English)

Text-to-audio (TTA) generation has made significant strides, yet achieving precise and consistent audio editing remains a major challenge. However, existing methods struggle to balance temporal consistency with background preservation. In this paper, we propose FreeSonic, a training-free framework leveraging the state-of-the-art Rectified Flow-based TangoFlux model. FreeSonic utilizes an optimized inversion-reverse process and joint text-audio attention maps for precise target segment extraction. For content editing, a novel scheduled attention decoupling confines modifications to target regions while preserving original acoustic context. Furthermore, task-oriented noise injection enhances versatility for tasks such as audio removal and non-rigid replacement. Extensive experimental results demonstrate that FreeSonic achieves a superior balance by providing a high-fidelity and efficient solution for precise and consistent audio editing. Project and demos: https://free-sonic.github.io/

音频编辑无训练注意力解耦TangoFlux

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。