arXiv:2506.11476cs.SDcs.LG2025-06中稿 · ISMIR 2025被引 5

轻量级控制网络让音乐生成可精细调控,内存占用更低

LiLAC: A Lightweight Latent ControlNet for Musical Audio Generation

  • 设计模块化轻量架构,参数量大幅减少
  • 音频质量与控制精度媲美传统ControlNet
  • 适合需要灵活、低资源部署的音乐生成场景

文本到音频的扩散模型能生成高质量且多样的音乐,但多数顶尖模型缺乏音乐制作所需的细粒度、随时间变化的控制能力。ControlNet通过克隆并微调编码器来附加外部控制,但会带来巨大的内存开销,并限制用户只能使用固定的控制集。本文提出一种轻量级、模块化的架构,在显著降低参数量的同时,保持与ControlNet相当的音频质量和条件遵循度。该方法提升了灵活性,大幅降低内存使用,使独立控制的训练与部署更加高效。我们进行了广泛的客观与主观评估,并在配套网站 https://lightlatentcontrol.github.io 提供了大量音频示例。

原文摘要 · Abstract (English)

Text-to-audio diffusion models produce high-quality and diverse music but many, if not most, of the SOTA models lack the fine-grained, time-varying controls essential for music production. ControlNet enables attaching external controls to a pre-trained generative model by cloning and fine-tuning its encoder on new conditionings. However, this approach incurs a large memory footprint and restricts users to a fixed set of controls. We propose a lightweight, modular architecture that considerably reduces parameter count while matching ControlNet in audio quality and condition adherence. Our method offers greater flexibility and significantly lower memory usage, enabling more efficient training and deployment of independent controls. We conduct extensive objective and subjective evaluations and provide numerous audio examples on the accompanying website at https://lightlatentcontrol.github.io

音乐生成控制网络轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。