用声音手势和文本生成高质量音频,可控且轻量。
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
- 通过音量、亮度、音高等时变信号控制音频生成
- 仅需40k步微调和单线性层,比ControlNet更轻量
- 支持语音模仿生成,兼顾文本语义与音频质量
我们提出Sketch2Sound,一种从可解释的时变控制信号(如音量、亮度、音高)和文本提示中生成高质量声音的生成模型。该模型可基于任意文本到音频的潜在扩散变压器(DiT),仅需40,000步微调及每控制信号一个线性层,显著轻于现有方法(如ControlNet)。为实现基于类似草图的声音模仿(如语音模仿或参考音形)的合成,训练中引入随机中值滤波,使控制信号具备灵活的时间粒度。实验表明,相较于纯文本基线,Sketch2Sound在遵循输入模仿的总体特征的同时,仍能保持文本语义一致性与音频质量。该模型让声音艺术家兼具文本的语义灵活性与声音手势的表达精度。音频示例见https://hugofloresgarcia.art/sketch2sound/。
原文摘要 · Abstract (English)
We present Sketch2Sound, a generative audio model capable of creating high-quality sounds from a set of interpretable time-varying control signals: loudness, brightness, and pitch, as well as text prompts. Sketch2Sound can synthesize arbitrary sounds from sonic imitations (i.e.,~a vocal imitation or a reference sound-shape). Sketch2Sound can be implemented on top of any text-to-audio latent diffusion transformer (DiT), and requires only 40k steps of fine-tuning and a single linear layer per control, making it more lightweight than existing methods like ControlNet. To synthesize from sketchlike sonic imitations, we propose applying random median filters to the control signals during training, allowing Sketch2Sound to be prompted using controls with flexible levels of temporal specificity. We show that Sketch2Sound can synthesize sounds that follow the gist of input controls from a vocal imitation while retaining the adherence to an input text prompt and audio quality compared to a text-only baseline. Sketch2Sound allows sound artists to create sounds with the semantic flexibility of text prompts and the expressivity and precision of a sonic gesture or vocal imitation. Sound examples are available at https://hugofloresgarcia.art/sketch2sound/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。