arXiv:2510.22950eess.AS2025-10被引 16

DiffRhythm 2 用块流匹配实现高效高保真歌曲生成,歌词与歌声对齐更精准。

DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching

  • 采用基于块流匹配的半自回归架构,无需外部标签即可对齐歌词与人声。
  • 音乐VAE帧率低至5 Hz,仍可实现高保真音频重建,提升长序列生成效率。
  • 提出交叉对偏好优化,支持多风格偏好训练,避免模型合并导致性能下降。

生成完整且高质量的歌曲极具挑战性,需在文本与音乐模态间以及音乐模态内部保持长期一致性。现有非自回归(NAR)框架虽能生成优质歌曲,但常出现歌词与人声对齐不佳的问题。同时,满足多样音乐偏好需依赖人类反馈强化学习(RLHF),但现有方法多通过合并多个模型进行多偏好优化,导致显著性能下降。为此,我们提出DiffRhythm 2,一种端到端的高保真可控歌曲生成框架。为解决歌词对齐问题,该框架采用基于块流匹配的半自回归架构,无需外部标签与约束即可实现歌词与演唱人声的精准对齐,同时保持NAR模型的高质量与高效率。为提升长序列生成的计算可行性,我们设计了音乐变分自编码器(VAE),实现5 Hz的低帧率,仍能保证高保真音频重建。此外,针对多偏好优化中模型合并带来的性能损失,我们提出交叉对偏好优化方法,有效缓解性能下降,支持在多样化人类偏好下更稳健的优化。我们还引入随机块表示对齐损失,进一步增强音乐表现力与结构连贯性。

原文摘要 · Abstract (English)

Generating full-length, high-quality songs is challenging, as it requires maintaining long-term coherence both across text and music modalities and within the music modality itself. Existing non-autoregressive (NAR) frameworks, while capable of producing high-quality songs, often struggle with the alignment between lyrics and vocal. Concurrently, catering to diverse musical preferences necessitates reinforcement learning from human feedback (RLHF). However, existing methods often rely on merging multiple models during multi-preference optimization, which results in significant performance degradation. To address these challenges, we introduce DiffRhythm 2, an end-to-end framework designed for high-fidelity, controllable song generation. To tackle the lyric alignment problem, DiffRhythm 2 employs a semi-autoregressive architecture based on block flow matching. This design enables faithful alignment of lyrics to singing vocals without relying on external labels and constraints, all while preserving the high generation quality and efficiency of NAR models. To make this framework computationally tractable for long sequences, we implement a music variational autoencoder (VAE) that achieves a low frame rate of 5 Hz while still enabling high-fidelity audio reconstruction. In addition, to overcome the limitations of multi-preference optimization in RLHF, we propose cross-pair preference optimization. This method effectively mitigates the performance drop typically associated with model merging, allowing for more robust optimization across diverse human preferences. We further enhance musicality and structural coherence by introducing stochastic block representation alignment loss.

歌曲生成扩散模型多偏好优化音乐合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。