用在线强化学习提升文本转音频质量,效果优于传统方法。
Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models
- 采用在线强化学习优化音频生成模型
- 仅470万参数即达新最佳性能
- 适合追求高音质与语义一致性的音频生成研究者
强化学习(RL)在大语言模型和视觉生成模型中已证明有效,但在文本到音频(TTA)生成领域仍少有探索。以往工作多采用离线方法如直接偏好优化(DPO),并利用对比语言-音频预训练(CLAP)模型作为奖励函数。本文研究将在线群组相对策略优化(GRPO)应用于TTA生成,适配基于流匹配的音频模型,结果表明在线RL显著优于离线方法。此外,引入大型音频语言模型(LALM)提供细粒度评分信号,更贴近人类感知。仅470万参数的最终模型Resonate,在TTA-Bench上实现了音频质量与语义对齐的新最佳表现。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has become an effective paradigm for enhancing Large Language Models (LLMs) and visual generative models. However, its application in text-to-audio (TTA) generation remains largely under-explored. Prior work typically employs offline methods like Direct Preference Optimization (DPO) and leverages Contrastive Language-Audio Pretraining (CLAP) models as reward functions. In this study, we investigate the integration of online Group Relative Policy Optimization (GRPO) into TTA generation. We adapt the algorithm for Flow Matching-based audio models and demonstrate that online RL significantly outperforms its offline counterparts. Furthermore, we incorporate rewards derived from Large Audio Language Models (LALMs), which can provide fine-grained scoring signals that are better aligned with human perception. With only 470M parameters, our final model, \textbf{Resonate}, establishes a new SOTA on TTA-Bench in terms of both audio quality and semantic alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。