用大模型生成细节音频描述,再用强化学习提升文本到音频生成质量
Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
- 先用大模型生成丰富音频描述,增强文本与音频语义对齐
- 采用GRPO算法微调扩散变换器,显著提升音质和提示遵循度
- 验证不同奖励函数效果,揭示音频生成中强化学习的关键设计
近年来,文本到音频(T2A)生成取得了显著进展,但现有方法在准确呈现复杂文本提示(尤其是包含复杂音频效果的提示)以及实现精确的文本-音频对齐方面仍面临挑战。尽管先前工作探索了数据增强、显式时间条件和强化学习等策略,整体合成质量仍受限。本文基于扩散变换器(DiT)架构,尝试通过强化学习进一步提升T2A生成质量。首先利用大语言模型(LLM)生成高保真、细节丰富的音频描述,显著改善模糊或信息不足提示下的文本-音频语义对齐。随后应用近期提出的组相对策略优化(GRPO)算法对T2A模型进行微调。通过系统性实验对比多种奖励函数(包括CLAP、KL、FAD及其组合),识别出影响音频合成效果的关键因素,并分析奖励设计对最终音频质量的影响。实验结果表明,基于GRPO的微调显著提升了合成保真度与提示遵循度。
原文摘要 · Abstract (English)
Text-to-audio (T2A) generation has advanced considerably in recent years, yet existing methods continue to face challenges in accurately rendering complex text prompts, particularly those involving intricate audio effects, and achieving precise text-audio alignment. While prior approaches have explored data augmentation, explicit timing conditioning, and reinforcement learning, overall synthesis quality remains constrained. In this work, we experiment with reinforcement learning to further enhance T2A generation quality, building on diffusion transformer (DiT)-based architectures. Our method first employs a large language model (LLM) to generate high-fidelity, richly detailed audio captions, substantially improving text-audio semantic alignment, especially for ambiguous or underspecified prompts. We then apply Group Relative Policy Optimization (GRPO), a recently introduced reinforcement learning algorithm, to fine-tune the T2A model. Through systematic experimentation with diverse reward functions (including CLAP, KL, FAD, and their combinations), we identify the key drivers of effective RL in audio synthesis and analyze how reward design impacts final audio quality. Experimental results demonstrate that GRPO-based fine-tuning yield substantial gains in synthesis fidelity and prompt adherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。