DiffRhythm+提升长篇歌曲生成质量与可控性,支持文本和音频风格控制。
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
- 基于扩散模型,融合文本与参考音频实现多模态风格控制。
- 在LJ-Speech、MUSDB18等数据集上生成更自然、结构更复杂的歌曲。
- 通过用户偏好优化,显著提升听感满意度与创作灵活性。
音乐作为人类智能与创造力的核心表现形式,近年来生成建模进展显著推动了长篇歌曲合成的发展。尽管先前的DiffRhythm模型已能生成带表达性人声与伴奏的完整歌曲,但仍受限于训练数据不平衡及风格控制不足,导致歌词重复或缺失、音乐质量不一致等问题。为此,我们提出DiffRhythm+,一个增强的基于扩散的框架,通过大幅扩充并平衡训练数据,有效缓解了上述问题,促进更丰富的音乐技能与表现力涌现。该框架引入多模态风格条件机制,允许用户通过描述性文本和参考音频精确指定音乐风格,显著提升创作控制力与多样性。此外,我们设计了直接对齐用户偏好的性能优化策略,在多项评估指标上引导模型生成更受青睐的结果。大量实验表明,DiffRhythm+在自然度、编排复杂度和听者满意度方面均显著优于现有系统。
原文摘要 · Abstract (English)
Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in long-form song generation, current systems for full-length song synthesis still face major challenges, including data imbalance, insufficient controllability, and inconsistent musical quality. DiffRhythm, a pioneering diffusion-based model, advanced the field by generating full-length songs with expressive vocals and accompaniment. However, its performance was constrained by an unbalanced model training dataset and limited controllability over musical style, resulting in noticeable quality disparities and restricted creative flexibility. To address these limitations, we propose DiffRhythm+, an enhanced diffusion-based framework for controllable and flexible full-length song generation. DiffRhythm+ leverages a substantially expanded and balanced training dataset to mitigate issues such as repetition and omission of lyrics, while also fostering the emergence of richer musical skills and expressiveness. The framework introduces a multi-modal style conditioning strategy, enabling users to precisely specify musical styles through both descriptive text and reference audio, thereby significantly enhancing creative control and diversity. We further introduce direct performance optimization aligned with user preferences, guiding the model toward consistently preferred outputs across evaluation metrics. Extensive experiments demonstrate that DiffRhythm+ achieves significant improvements in naturalness, arrangement complexity, and listener satisfaction over previous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。