arXiv:2605.03821cs.ROcs.AI2026-05被引 2

用奖励对齐提升机器人视频模型的任务一致性与真实感

RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models

论文配图:RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models
图 1 · 摘自论文原文
  • 用多模态裁判模型评估生成视频,蒸馏为轻量奖励模型用于强化学习后训练
  • 在10,000对视频-指令数据上提升六维评分10.1%,操纵准确率增7.5%
  • 引入滑动窗口重编码,仅增加1%延迟却显著改善长序列预测质量

现有机器人视频世界模型通常采用重建和感知相似性等低层目标进行训练,这些目标与机器人决策最关键的性能——如指令遵循、操作成功率和物理合理性——严重不匹配,且在长时序自回归预测中易积累误差。本文提出RoboAlign-R1框架,结合奖励对齐的后训练与稳定的长时序推理。构建了包含10,000个标注视频-指令对的RobotWorldBench基准,训练多模态教师裁判RoboAlign-Judge,实现生成视频的六维细粒度评估。随后将教师模型蒸馏为轻量学生奖励模型,支持高效强化学习后训练。为缓解长时序生成漂移,提出无需训练的滑动窗口重编码(SWR)策略,定期刷新生成上下文。在域内评估下,RoboAlign-R1相比最强基线综合六维得分提升10.1%,其中操作准确率提升7.5%,指令遵循提升4.6%;外部VLM交叉验证与盲测人类评估均支持此排名提升。同时,SWR仅带来约1%额外延迟,却使SSIM提升2.8%,LPIPS降低9.8%。结果表明,奖励对齐后训练与稳定长时推理可有效提升机器人视频世界模型的任务一致性、物理真实性与长时预测质量。

原文摘要 · Abstract (English)

Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities that matter most for robot decision making, including instruction following, manipulation success, and physical plausibility. They also suffer from error accumulation in long-horizon autoregressive prediction. We present RoboAlign-R1, a framework that combines reward-aligned post-training with stabilized long-horizon inference for robot video world models. We construct RobotWorldBench, a benchmark of 10,000 annotated video-instruction pairs collected from four robot data sources, and train a multimodal teacher judge, RoboAlign-Judge, to provide fine-grained six-dimensional evaluation of generated videos. We then distill the teacher into a lightweight student reward model for efficient reinforcement-learning-based post-training. To reduce long-horizon rollout drift, we further introduce Sliding Window Re-encoding (SWR), a training-free inference strategy that periodically refreshes the generation context. Under our in-domain evaluation protocol, RoboAlign-R1 improves the aggregate six-dimension score by 10.1% over the strongest baseline, including gains of 7.5% on Manipulation Accuracy and 4.6% on Instruction Following; these ranking improvements are further supported by an external VLM-based cross-check and a blinded human study. Meanwhile, SWR improves long-horizon prediction quality with only about 1% additional latency, yielding a 2.8% gain in SSIM and a 9.8% reduction in LPIPS. Together, these results show that reward-aligned post-training and stabilized long-horizon decoding improve task consistency, physical realism, and long-horizon prediction quality in robot video world models.

机器人视频生成奖励对齐长时序预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。