让世界模型更懂动作如何改变世界,提升生成视频的物理一致性。
CAER: Causal Action Effect Reweighting for World Model Training

- 通过对比有无动作条件的预测,动态定位受动作影响的帧
- 训练时重点优化因果相关区域,视觉质量与可控性显著提升
- 无需额外标注,可直接用于各类动作条件世界模型
世界模型正成为具身智能的核心基础设施,动作条件视频生成可实现对环境演变的可控预测。然而现有模型通常采用时空均匀的均方误差训练,导致大量背景像素主导梯度,而稀疏的动作交互动态被忽略;这种均匀拟合更倾向于复现外观而非学习动作如何改变世界。本文提出因果动作效应重加权(CAER),一种通用训练范式,将监督信号重新分配至受动作因果影响的预测帧上。CAER在线比较有无动作条件的预测结果,定位受影响区域,生成权重图并归一化,保持总权重不变,仅调整分配位置。该信号无需外部标注或离线预处理,不增加数据处理开销,且随模型和数据规模自然扩展。跨多种异构动作条件世界模型任务的实验表明,相较于均匀均方误差训练,CAER能更快收敛至更优解,在物理一致性、可控性和视觉质量方面均有稳定提升。
原文摘要 · Abstract (English)
World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agent interventions. Yet existing models are commonly trained with space-time-uniform mean squared error, allowing abundant background tokens to dominate the gradient while sparse interaction dynamics remain under-optimized; such uniform fitting rewards reconstructing appearance rather than learning how actions change the world. We introduce Causal Action Effect Reweighting (CAER), a general training paradigm that redistributes supervision toward the tokens whose predicted future is causally affected by the action. CAER contrasts the model's own predictions with and without action conditioning to localize these tokens online, then normalizes the resulting effect map into a weight that preserves the total coefficient mass and changes only where it is spent. This online signal requires no external annotations or offline preprocessing, avoids additional data-processing time, and scales naturally with model and dataset size. Experiments across heterogeneous action-conditioned world-model tasks show that CAER converges to better solutions than uniform MSE training, with consistent improvements in the physical consistency, controllability, and visual quality of generated videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。