arXiv:2609.03952cs.CV2026-09

用视觉语言模型统一评估视频动作一致性和视觉质量。

WorldReward: Reward Modeling for Camera-Conditioned World Models

论文配图:WorldReward: Reward Modeling for Camera-Conditioned World Models
图 1 · 摘自论文原文
  • 将视频按动作对齐分块,结构化提取视觉证据
  • 在三个维度上超越GPT-5.5,提升3.42~3.56个百分点
  • 适合需要精准控制动作与画面质量的视频生成任务

相机条件世界模型需生成交互式视频,使动作指令引发预期场景变化,同时保持外观、几何和时间动态的一致性。现有奖励机制分别评估:基于几何的奖励可衡量轨迹执行但无法判断视觉质量;基于图像的奖励能评估帧质量却忽略动作执行与时间动态。本文提出WorldReward,一种基于视觉语言模型(VLM)的成对偏好奖励模型,统一评估动作一致性与视觉质量。它将成对视频分解为动作对齐的片段,结构化组织每段视觉证据,并通过投票聚合得到视频级的动作与视觉质量偏好。训练采用大规模推理增强的偏好数据集,由前沿VLM生成结构化判断,经工具代理审计与定向人工审查优化。进一步构建WorldReward-Bench,一个包含人类标注的基准,评估奖励模型与人类偏好在动作一致性、外观质量、运动质量上的匹配度。结果表明,WorldReward在三方面均取得最高一致性,分别超过GPT-5.5 3.42、1.45、3.56个百分点。在对HY-WorldPlay 1.5进行强化学习后训练时,显著提升短至长时程下的动作执行与视觉质量。

原文摘要 · Abstract (English)

Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.

视频生成视觉语言模型奖励建模世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。