arXiv:2608.24885cs.ROcs.CV2026-08

提升机器人世界模型对任意动作的生成准确性,让模拟更可靠。

Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning

论文配图:Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning
图 1 · 摘自论文原文
  • 用新基准测试模型对非专家动作的跟随能力
  • 模型对非专家动作常忽略或生成无效轨迹
  • 提出三重改进方法,适合强化学习与真实机器人应用

动作条件世界模型被广泛用于策略评估与优化,但其有效性依赖于一个未经验证的假设:生成的未来能准确反映任意有效动作。现有基准多局限于专家示范,对非专家动作的评估不足。为此,我们提出WorldEcho,通过视觉完整性与SE(3)轨迹对齐,在更广动作分布上诊断动作跟随能力。诊断显示,当前模型虽能合理执行专家动作,但在多样非专家轨迹上表现不佳,或忽略指令,或生成视觉无效结果。我们进一步提出WorldSync,从分布覆盖、表征基础和干预效果对齐三个维度增强动作跟随:扩大动作后果训练分布,通过动作强制专家将中间视频表示与动作引发的机器人动力学对齐,并使预测变化与真实未来变化一致。在RoboTwin基准和真实机器人任务上的实验表明,WorldSync显著提升WorldEcho指标,作为迭代策略改进的更可靠模拟器,使策略成功率更高。

原文摘要 · Abstract (English)

Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithfully reflect arbitrary valid actions. Existing benchmarks are typically confined to expert demonstrations, leaving off-expert action following inadequately evaluated. To address this gap, we introduce WorldEcho, which probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. Our diagnosis shows that current world models reasonably execute expert actions but struggle with diverse off-expert trajectories, either ignoring the commanded actions or producing visually invalid rollouts. We further propose WorldSync, which strengthens action following along three complementary axes: distributional coverage, representational grounding, and intervention-effect alignment. It broadens the training distribution over action consequences, grounds intermediate video representations in action-induced robot dynamics through an Action-Forcing Expert, and aligns predicted changes under action interventions with the corresponding changes in ground-truth futures. Experiments on RoboTwin benchmarks and real-robot tasks show that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to achieve higher success rates.

世界模型动作跟随机器人仿真策略学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。