arXiv:2510.12175cs.SDeess.AS2025-10中稿 · publication in the…被引 1

用四种声学信号精准控制音效生成,保持高质量与文本对齐。

Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis

  • 引入响度、音高、频谱质心和音色四类动态控制信号。
  • 仅需原模型0.85%参数即可实现高质量音效合成。
  • 适合音乐创作与音效设计领域的艺术家和研究者使用。

基于扩散模型的生成技术虽已实现高质量文本到音频合成,但细粒度声学控制仍是开源研究中的难题。本文提出Audio Palette,一种基于扩散Transformer(DiT)的模型,扩展了Stable Audio Open架构以解决这一‘控制差距’。不同于仅依赖语义条件的方法,Audio Palette引入响度、音高、频谱质心和音色四类时变控制信号,实现对声学特征的精确且可解释的调控。模型通过在精选的AudioSet子集上使用低秩适配(LoRA)高效适应福莱音效领域,仅需训练原模型0.85%的参数。实验表明,Audio Palette在保持高音频质量与文本提示强对齐的同时,实现了细粒度可解释的声学属性控制,其在标准指标如弗雷切特音频距离(FAD)和LAION-CLAP得分上与基线模型相当。本文提供了一套可扩展、模块化的音频研究流程,强调序列化条件输入、内存效率及三尺度无分类器引导机制,支持推断时的精细控制。该工作为开源环境下的可控声音设计与表演性音频合成奠定了坚实基础,推动音乐与声音信息检索领域的艺术家导向型工作流发展。

原文摘要 · Abstract (English)

Recent advances in diffusion-based generative models have enabled high-quality text-to-audio synthesis, but fine-grained acoustic control remains a significant challenge in open-source research. We present Audio Palette, a diffusion transformer (DiT) based model that extends the Stable Audio Open architecture to address this "control gap" in controllable audio generation. Unlike prior approaches that rely solely on semantic conditioning, Audio Palette introduces four time-varying control signals, loudness, pitch, spectral centroid, and timbre, for precise and interpretable manipulation of acoustic features. The model is efficiently adapted for the nuanced domain of Foley synthesis using Low-Rank Adaptation (LoRA) on a curated subset of AudioSet, requiring only 0.85% of the original parameters to be trained. Experiments demonstrate that Audio Palette achieves fine-grained, interpretable control of sound attributes. Crucially, it accomplishes this novel controllability while maintaining high audio quality and strong semantic alignment to text prompts, with performance on standard metrics such as Frechet Audio Distance (FAD) and LAION-CLAP scores remaining comparable to the original baseline model. We provide a scalable, modular pipeline for audio research, emphasizing sequence-based conditioning, memory efficiency, and a three-scale classifier-free guidance mechanism for nuanced inference-time control. This work establishes a robust foundation for controllable sound design and performative audio synthesis in open-source settings, enabling a more artist-centric workflow in the broader context of music and sound information retrieval.

音效生成扩散模型可控合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。