arXiv:2509.05256cs.SDcs.AI2025-09被引 6

用文本和时间轴编辑复杂声音场景中的单个音效,支持删、插、增强。

Recomposer: Event-roll-guided generative audio editing

  • 基于事件轨时间图与文本描述,通过Transformer模型精准定位音效位置。
  • 在真实背景上合成音效对训练,实现删除、插入、增强音效的高保真还原。
  • 适合音频制作、影视后期等需要精细控制声音内容的场景。

在复杂真实声景中编辑单个声音事件十分困难,因声音源在时间上重叠。生成模型可基于其对数据域的强先验知识,补全缺失或受损的细节。我们提出一种系统,可基于文本编辑描述(如“增强门声”)和由“事件轨”转录得到的事件时间图形表示,对复杂场景中的单个声音事件进行删除、插入和增强。该系统采用基于SoundStream表征的编码器-解码器变压器架构,在合成的(输入,期望输出)音频对上训练,这些对通过将孤立音效叠加到密集的真实背景上生成。评估表明,编辑描述中的操作、类别和时间信息均至关重要。本工作证明了‘重组’是一项重要且实用的应用。

原文摘要 · Abstract (English)

Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their strong prior understanding of the data domain. We present a system for editing individual sound events within complex scenes able to delete, insert, and enhance individual sound events based on textual edit descriptions (e.g., ``enhance Door'') and a graphical representation of the event timing derived from an ``event roll'' transcription. We present an encoder-decoder transformer working on SoundStream representations, trained on synthetic (input, desired output) audio example pairs formed by adding isolated sound events to dense, real-world backgrounds. Evaluation reveals the importance of each part of the edit descriptions -- action, class, timing. Our work demonstrates ``recomposition'' is an important and practical application.

音频编辑生成模型事件轨SoundStream

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。