arXiv:2606.23139eess.AS2026-06综述被引 2

梳理基础模型时代的音频编辑任务与主流方法,助力AIGC音频创作。

Audio Editing in the Era of Foundation Models: A Survey

论文配图:Audio Editing in the Era of Foundation Models: A Survey
图 1 · 摘自论文原文
  • 构建统一任务分类体系,涵盖多种音频编辑需求
  • 总结训练型与免训练型基础模型的代表性方法
  • 适合从事音频生成、AIGC研究者参考

音频编辑旨在修改给定的合成或真实音频信号以满足特定用户需求。作为AIGC中一个有前景但具挑战性的方向,近年来受到越来越多关注。音频生成技术的进展使强大的生成模型成为现代音频编辑系统的核心。这一快速发展催生了对新兴任务、方法和资源进行系统梳理的迫切需求。本文全面回顾了基础模型时代的音频编辑研究。首先提出一个统一的编辑任务分类体系,随后总结支持现代音频编辑的主要基础模型范式,涵盖训练型与免训练型代表性方法。进一步讨论相关资源,包括数据集、评估协议和数据构建工具。最后,识别该领域的开放挑战,并展望未来研究方向。项目主页已发布于 https://github.com/DaViD-Pigeon/AudioEditSurvey。

原文摘要 · Abstract (English)

Audio editing aims to modify a given synthetic or real-world audio signal to satisfy specific user needs. As a promising yet challenging direction in AIGC, it has attracted increasing attention. Recent advances in audio generation have made powerful generative models central to modern audio editing systems. This rapid progress has created a growing need to organize emerging tasks, methods, and resources into a coherent view. In this survey, we provide a comprehensive review of audio editing in the era of foundation models. We first present a unified taxonomy of existing editing tasks and then summarize the major foundation-model paradigms that support modern audio editing, covering representative approaches from both training-based and training-free perspectives. We further discuss related resources, including datasets, evaluation protocols, and data construction tools. Finally, we identify open challenges in this field and outline promising directions for future research. The project page is released at https://github.com/DaViD-Pigeon/AudioEditSurvey.

音频编辑AIGC基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。