arXiv:2608.04653cs.CVcs.RO2026-08

提出新框架CoCo,让世界模型真正响应动作而非依赖视觉惯性。

Overcoming Statistical Bias in Action-Controllable World Models

论文配图:Overcoming Statistical Bias in Action-Controllable World Models
图 1 · 摘自论文原文
  • 通过反事实一致性约束,强制模型在动作变化时预测结果一致。
  • 在Mini-SSMB上动作响应一致性提升至0.483,漂移能量降低17.07%。
  • 适合需要精准动作控制的机器人规划与视觉预测任务。

动作条件世界模型旨在预测智能体动作下视觉环境的演化。然而,未来帧常仅靠视觉惯性和重复运动模式即可高度预测,导致模型依赖统计偏差而非真实动作依赖,使不同动作产生相似未来,甚至零动作下仍持续运动。为此,本文提出反事实一致性框架CoCo,包含多步反事实一致性与动作-空间反事实一致性,通过约束参考、逆动作与零动作轨迹,以及镜像场景与动作变换下的预测一致性,削弱对统计捷径的依赖。引入动作响应一致性(ARC)与漂移能量(DE)评估可控性,并设计Mini-SSMB用于同状态多动作反事实评估。在Mini-SSMB上,模型达到ARC_inv 0.412、ARC_ref 0.483,DE相对基线降低17.07%;在VP2视觉规划中,成功率达73.1%,优于现有最优模型。在BAIR与RoboNet数据集上验证,性能提升同时保持视频预测质量并具备跨模型迁移能力。

原文摘要 · Abstract (English)

Action-conditioned world models aim to predict how visual environments evolve under an agent's actions. Yet future frames are often highly predictable from visual inertia and recurring motion patterns alone. This creates a shortcut: models can fit the data by exploiting statistical biases without making their visible dynamics meaningfully depend on the action. As a result, different actions may produce similar futures, while motion may persist even under zero action. The key question is how to reduce reliance on statistical shortcuts from dominating action-conditioned prediction. We argue that action control requires more than injecting action features; it requires enforcing consistency under counterfactual changes to actions and observations. Based on this insight, we introduce CoCo, a Counterfactual Consistency framework to enhance action controllability through two complementary constraints. Multi-step counterfactual consistency constrains reference, inverse-action, and zero-action rollouts, while action-spatial counterfactual consistency enforces consistent predictions under mirrored scenes and transformed actions. Together, they reduce reliance on statistical shortcuts from substituting for action-dependent dynamics. We further introduce Action Response Consistency (ARC) and Drift Energy (DE) to assess action controllability, together with Mini-SSMB for same-state, multi-action counterfactual evaluation. On Mini-SSMB, our full model achieved ARC_inv of 0.412 and ARC_ref of 0.483, while reducing DE by 17.07% relative to the baseline. On VP2 visual planning, it achieves the highest average success rate among SOTA models, at 73.1%. Experiments on BAIR and RoboNet further show that these gains preserve video prediction quality and transfer across model settings.

世界模型动作控制反事实学习视觉预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。