用人类偏好训练联合音视频生成的奖励模型,避免只看指标得分却忽略整体协调性。
VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

- 基于人类对比反馈构建链式思维奖励模型,分步提升判断准确性。
- 在10.3万组对比中表现优于传统指标,对齐真实人类偏好。
- 适合做音视频生成质量评估或后训练优化的研究者使用。
利用强化学习对联合音视频生成模型进行后训练需要奖励信号。现有方法通过组合音频质量、视觉保真度和同步性等独立指标构建奖励,但这些指标分别评估感知维度,无法捕捉文本提示、视频与音频之间整体语义与时间上的连贯性,导致模型为追求高分而出现不连贯或不符合人类直觉的内容。为此,我们构建了大规模人类偏好数据集 VAPref-10K,包含9,000个提示和10.3万组细粒度成对比较,来自开源生成模型。同时引入 VA-Judger-Bench 基准,涵盖域内与域外模型对比,以检验奖励模型是否真正符合人类偏好。我们提出 VA-Judger,一种链式思维的全维奖励模型:先从差异明显的对比中学习结构化输出与粗粒度偏好判断,再通过拒绝采样验证人类标注,提炼困难样本的可靠解释,最后进行维度分解的强化学习,将人类反馈拆解为各质量维度,获得比单一二分类标签更密集的奖励信号。实验表明,VA-Judger 在域内与域外评估中均优于基于指标的基线,在预测人类偏好上表现更优;使用其对齐人类偏好的奖励信号进行后训练,显著提升了音视频生成质量。
原文摘要 · Abstract (English)
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large-scale human-preference dataset VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from open-source generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. In particular, VA-Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards for post-training audio-video generation model also yields significant improvements in generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。