Stable Audio 3可快速生成多分钟音频并支持精准编辑,适合音乐与声音创作。
Stable Audio 3
- 采用新型语义-声学自编码器,将音频压缩至紧凑隐空间进行扩散生成。
- 支持变量长度生成与音频修复,2秒内完成短音频生成,推理步数更少。
- 模型轻量高效,小/中版本可在消费级硬件运行,适合创作者快速实验。
Stable Audio 3 是一组用于可变长度音频生成与编辑的快速潜在扩散模型(小型、中型、大型)。由于模型可生成数分钟音频,可变长度生成能避免为短音频生成完整时长带来的成本。同时支持图像修复(inpainting),实现音频目标编辑及短录音续写。其潜在扩散模型基于一种新型语义-声学自编码器,将音频映射到紧凑隐空间,在保证音频保真度的同时增强隐空间的语义结构。此外,通过对抗性后训练加速推理并提升生成质量,减少推理步数,改善音质与提示遵循度。Stable Audio 3 模型使用授权及 Creative Commons 数据训练,可在 H200 GPU 上实现小于 2 秒的音频生成,MacBook Pro M4 上为数秒内完成。我们发布了小型和中型模型权重及其训练与推理流程,可在消费级硬件上运行。
原文摘要 · Abstract (English)
Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of audio, variable-length generations are key to avoid the cost of producing full-length generations for short sounds. We also support inpainting, enabling targeted audio editing and the continuation of short recordings. Our latent diffusion models operate on top of a novel semantic-acoustic autoencoder that projects audio into a compact latent space, enabling efficient diffusion-based generation while preserving audio fidelity and encouraging semantic structure in the latent. Finally, we run adversarial post-training to both accelerate inference and improve generation quality, reducing the number of inference steps while improving fidelity and prompt adherence. Stable Audio 3 models are trained on licensed and Creative Commons data to generate music and sounds in less than a 2s on an H200 GPU and less than a few seconds on a MacBook Pro M4. We release the weights of small and medium, that can run on consumer-grade hardware, together with their training and inference pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。