arXiv:2504.09219cs.SDeess.AS2025-04

用文字生成音乐音色,让创作者自定义电子乐器声音。

Generation of Musical Timbres using a Text-Guided Diffusion Model

  • 结合扩散模型与多模态对比学习,直接从文本生成音色频谱。
  • 联合生成频谱幅度与相位,无需后续相位恢复步骤。
  • 适合音乐人、作曲家快速构建个性化电子音源。

近年来,文本到音频系统取得了显著进展,能够直接从文本描述生成完整的音频片段。尽管这些系统也支持音乐创作,但人类的创造力和表达意图常被限制。本文提出一种新方法,使作曲家、编曲者和演奏者能创建音乐创作的基本单元:用于电子乐器和数字音频工作站(DAWs)的单个乐音音频。用户通过文本提示指定音色特征。我们引入一个结合潜在扩散模型与多模态对比学习的系统,实现基于文本描述的音乐音色生成。该方法联合生成频谱的幅度与相位,避免了如以往方法需额外运行相位恢复算法的问题。音频示例、源代码及网页应用已公开:https://wxuanyuan.github.io/Musical-Note-Generation/

原文摘要 · Abstract (English)

In recent years, text-to-audio systems have achieved remarkable success, enabling the generation of complete audio segments directly from text descriptions. While these systems also facilitate music creation, the element of human creativity and deliberate expression is often limited. In contrast, the present work allows composers, arrangers, and performers to create the basic building blocks for music creation: audio of individual musical notes for use in electronic instruments and DAWs. Through text prompts, the user can specify the timbre characteristics of the audio. We introduce a system that combines a latent diffusion model and multi-modal contrastive learning to generate musical timbres conditioned on text descriptions. By jointly generating the magnitude and phase of the spectrogram, our method eliminates the need for subsequently running a phase retrieval algorithm, as related methods do. Audio examples, source code, and a web app are available at https://wxuanyuan.github.io/Musical-Note-Generation/

音色生成扩散模型文本生成音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。