用细粒度引导提升扩散模型作曲精度与可听性
Efficient Fine-Grained Guidance for Diffusion Model Based Symbolic Music Generation
- 在扩散模型中引入细粒度引导,精准控制音高生成
- 实测与主观评价显示音乐质量显著提升,支持即兴创作
- 适合音乐生成、交互式作曲等高级应用
由于符号化音乐生成面临数据稀缺和音高精度要求高的双重挑战,本文提出一种高效的细粒度引导(FGG)方法,用于改进扩散模型的生成能力。该方法使模型更贴近专业作曲者的控制意图,显著提升生成音乐的准确性、可听性与整体质量,适用于即兴演奏和交互式音乐创作等复杂场景。我们给出了理论分析,并通过数值实验与主观评测验证了方法的有效性。相关演示页面已发布,支持实时交互生成。
原文摘要 · Abstract (English)
Developing generative models to create or conditionally create symbolic music presents unique challenges due to the combination of limited data availability and the need for high precision in note pitch. To address these challenges, we introduce an efficient Fine-Grained Guidance (FGG) approach within diffusion models. FGG guides the diffusion models to generate music that aligns more closely with the control and intent of expert composers, which is critical to improve the accuracy, listenability, and quality of generated music. This approach empowers diffusion models to excel in advanced applications such as improvisation, and interactive music creation. We derive theoretical characterizations for both the challenges in symbolic music generation and the effects of the FGG approach. We provide numerical experiments and subjective evaluation to demonstrate the effectiveness of our approach. We have published a demo page to showcase performances, which enables real-time interactive generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。