arXiv:2608.21425cs.CVcs.AI2026-08中稿 · ECCV

用分布对齐提升视频生成的人类偏好一致性

Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation

  • 通过精英筛选校准人类偏好数据,提升奖励信号可靠性
  • 将视频质量建模为多维奖励分布,捕捉人类偏好的不确定性
  • 用Wasserstein距离对齐分布,更好匹配人类偏好的全局结构

视频生成是人工智能内容创作的核心。对齐生成视频与人类偏好是评估生成质量的关键标准。尽管视觉质量已有显著进展,仍存在三大挑战:其一,奖励信号可靠性受限于人类偏好数据的质量,常受主观噪声和偏差影响;其二,标准标量奖励模型将多维度人类偏好压缩为单一数值,导致不同偏好维度间动态权衡的丢失;其三,在策略优化中,广泛使用的KL散度仅施加局部约束,难以捕捉人类偏好的全局结构。为此,我们提出统一的、面向偏好的视频生成学习框架。首先引入精英引导过滤,校准偏好数据并构建可靠监督信号用于奖励模型训练。随后,将视频质量建模为多维奖励分布,以捕捉人类偏好的固有不确定性,并使用Wasserstein距离对齐学习到的奖励分布与实际人类偏好分布。最后,将基于Wasserstein的分布对齐引入GRPO,引导策略优化更准确地匹配人类偏好的全局结构。在奖励建模与视频生成任务上的实验表明,该方法显著提升了奖励信号的可靠性与生成视频的感知一致性。代码已开源:https://github.com/alignhs26/ahs。

原文摘要 · Abstract (English)

Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluating generation quality. Despite significant progress in visual quality, three key challenges remain. First, the reliability of reward signals is constrained by the quality of human preference data, which is often affected by subjective noise and bias. Second, standard scalar reward models collapse multi-aspect human preferences into a single value, leading to the loss of dynamic trade-offs across multiple preference dimensions. Third, in policy optimization, the widely adopted KL divergence imposes primarily local constraints and may fail to capture the global structure of human preferences. To address these challenges, we propose a unified preference-aware learning framework for video generation. First, we introduce elite-guided filtering to calibrate preference data and construct reliable supervision for reward model training. We then model video quality as a multidimensional reward distribution to capture the uncertainty inherent in human preferences, and use the Wasserstein distance to align the learned reward distribution with the empirical human preference distribution. Finally, we introduce Wasserstein-based distributional alignment into GRPO, guiding policy optimization to better match the global structure of human preferences over videos. Experiments on reward modeling and video generation demonstrate that our approach improves the reliability of reward signals and the perceptual consistency of generated videos. Our code is available at https://github.com/alignhs26/ahs.

视频生成偏好对齐分布对齐奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。