arXiv:2505.16076eess.AS2025-05被引 12

无需训练即可精准编辑音频,保持原音质量

AudioMorphix: Training-free audio editing with diffusion probabilistic models

  • 直接操作频谱图,通过参考音频实现局部精准修改
  • 支持添加、移除、变调等操作,保持高保真与精确度
  • 适合音视频创作者、音频修复人员使用

精准编辑声音是音频内容创作中的关键挑战,现有方法多依赖文本指令或音频样本对,难以在保持原始录音保真度的同时实现精细修改。本文提出AudioMorphix,一种无需训练的音频编辑方法,通过直接操作频谱图,在指定时频区域进行局部修改,其余部分保持不变。受形态学理论启发,将音频混合视为可逆的渐变过程,利用能量函数引导扩散过程,优化噪声潜在变量以融合输入与参考音频特征。同时引入缓存机制增强自注意力层,保留原始录音的细节特征。为推动研究,我们构建了包含多样化编辑指令的新评估基准。大量实验表明,AudioMorphix在添加、移除、时间偏移与拉伸、变调等多种任务中表现优异,兼具高保真与精确性。

原文摘要 · Abstract (English)

Editing sound with precision is a crucial yet underexplored challenge in audio content creation. While existing works can manipulate sounds by text instructions or audio exemplar pairs, they often struggled to modify audio content precisely while preserving fidelity to the original recording. In this work, we introduce a novel editing approach that enables localized modifications to specific time-frequency regions while keeping the remaining of the audio intact by operating on spectrograms directly. To achieve this, we propose AudioMorphix, a training-free audio editor that manipulates a target region on the spectrogram by referring to another recording. Inspired by morphing theory, we conceptualize audio mixing as a process where different sounds blend seamlessly through morphing and can be decomposed back into individual components via demorphing. Our AudioMorphix optimizes the noised latent conditioned on raw input and reference audio while rectifying the guided diffusion process through a series of energy functions. Additionally, we enhance self-attention layers with a cache mechanism to preserve detailed characteristics from the original recordings. To advance audio editing research, we devise a new evaluation benchmark, which includes a curated dataset with a variety of editing instructions. Extensive experiments demonstrate that AudioMorphix yields promising performance on various audio editing tasks, including addition, removal, time shifting and stretching, and pitch shifting, achieving high fidelity and precision. Demo and code are available at this url.

音频编辑扩散模型无训练频谱操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。