用标准化奖励提升文本到音频生成的逼真度与对齐度
SCORE: Scaling audio generation using Standardized COmposite REwards
- 引入推理时缩放,通过加权融合多维度奖励提升生成质量
- 在AudioCaps和MUSDB18上实现显著更高的语义对齐与听觉质量
- 适合需要精准控制音频生成效果的研究者与开发者
本文旨在提升文本到音频生成在推理阶段的表现,生成与文本提示高度一致且逼真的音频。尽管进展迅速,现有模型常难以平衡感知质量与文本对齐。为此,我们采用无需训练的推理时缩放方法,增加推理计算以提升性能,并首次将其应用于音频生成。提出一种新型多奖励引导机制,将感知、语义等关键成分的奖励值标准化至统一尺度后加权求和,确保引导稳定并支持显式控制。此外,引入基于音频语言模型的新对齐评估指标,实现更鲁棒的评价。实验表明,该方法在音频-文本对齐与感知质量上均显著优于基线及现有奖励引导技术。合成样例可在项目主页查看:https://mm.kaist.ac.kr/projects/score
原文摘要 · Abstract (English)
The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advancements, existing models often fail to achieve a reliable balance between perceptual quality and textual alignment. To address this, we adopt Inference-Time Scaling, a training-free method that improves performance by increasing inference computation. We establish its unexplored application to audio generation and propose a novel multi-reward guidance that equally signifies each component essential in perception. By normalizing each reward value into a common scale and combining them with a weighted summation, the method not only enforces stable guidance but also enables explicit control to reach desired aspects. Moreover, we introduce a new audio-text alignment metric using an audio language model for more robust evaluation. Empirically, our method improves both semantic alignment and perceptual quality, significantly outperforming naive generation and existing reward guidance techniques. Synthesized samples are available on our demo page: https://mm.kaist.ac.kr/projects/score
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。