用扩散Transformer提升伴奏生成质量与速度,支持文本控制。
Improving Musical Accompaniment Co-creation via Diffusion Transformers
- 用扩散Transformer替代原有模型,提升音质与多样性。
- 通过跨模态预测网络实现文本到音频的精准映射。
- 减少去噪步骤仍保持高质量,推理速度更快,适合音乐创作应用。
在基于潜在扩散模型的伴奏生成框架Diff-A-Riff基础上,我们提出多项改进:将底层自编码器升级为支持立体声、保真度更高的模型,并用扩散Transformer替换原有的潜空间U-Net;通过训练跨模态预测网络,将文本提取的CLAP嵌入转换为音频对应的CLAP嵌入,优化文本提示效果;最后采用一致性训练框架加速推理,在显著减少去噪步骤的同时保持优异生成质量。通过消融实验和客观指标评估,新模型在生成质量、多样性、推理速度和文本控制能力上均优于原版Diff-A-Riff。音频样例可访问:https://sonycslparis.github.io/improved_dar/
原文摘要 · Abstract (English)
Building upon Diff-A-Riff, a latent diffusion model for musical instrument accompaniment generation, we present a series of improvements targeting quality, diversity, inference speed, and text-driven control. First, we upgrade the underlying autoencoder to a stereo-capable model with superior fidelity and replace the latent U-Net with a Diffusion Transformer. Additionally, we refine text prompting by training a cross-modality predictive network to translate text-derived CLAP embeddings to audio-derived CLAP embeddings. Finally, we improve inference speed by training the latent model using a consistency framework, achieving competitive quality with fewer denoising steps. Our model is evaluated against the original Diff-A-Riff variant using objective metrics in ablation experiments, demonstrating promising advancements in all targeted areas. Sound examples are available at: https://sonycslparis.github.io/improved_dar/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。