用轻量ControlNet实现音频生成的精细调控,支持音量、音高和事件编辑。
Audio ControlNet for Fine-Grained Audio Generation and Editing
- 在预训练模型上加ControlNet,无需重训即可控制音量、音高和事件
- 仅38M额外参数,音频集强监督下事件与段落级F1达当前最优
- 可按指令精准删减或插入音频事件,适合音效设计与创意编辑
我们研究细粒度文本到音频(T2A)生成任务。尽管现有模型能从文本生成高质量音频,但对音量、音高和声音事件等属性的控制能力不足。不同于需为每种控制类型重新训练模型的旧方法,我们提出在预训练T2A主干上训练ControlNet模型,实现对音量、音高和事件序列的可控生成。提出两种设计:T2A-ControlNet与T2A-Adapter,其中T2A-Adapter结构更高效且控制能力强。仅引入3800万额外参数,T2A-Adapter在AudioSet-Strong数据集上事件级与段落级F1分数均达到当前最佳。进一步扩展该框架至音频编辑,提出T2A-Editor,可根据指令在指定时间位置移除或插入音频事件。模型、代码、数据管道与评测基准将公开,以支持可控音频生成与编辑的后续研究。
原文摘要 · Abstract (English)
We study the fine-grained text-to-audio (T2A) generation task. While recent models can synthesize high-quality audio from text descriptions, they often lack precise control over attributes such as loudness, pitch, and sound events. Unlike prior approaches that retrain models for specific control types, we propose to train ControlNet models on top of pre-trained T2A backbones to achieve controllable generation over loudness, pitch, and event roll. We introduce two designs, T2A-ControlNet and T2A-Adapter, and show that the T2A-Adapter model offers a more efficient structure with strong control ability. With only 38M additional parameters, T2A-Adapter achieves state-of-the-art performance on the AudioSet-Strong in both event-level and segment-level F1 scores. We further extend this framework to audio editing, proposing T2A-Editor for removing and inserting audio events at time locations specified by instructions. Models, code, dataset pipelines, and benchmarks will be released to support future research on controllable audio generation and editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。