用群作用理论评估视频世界模型的动作真实性,提升动态合理性。
World Models as Group Actions

- 将动作结构建模为群作用,通过隐空间正则化实现一致性约束。
- 在多个先进模型上提升群作用一致性和滚动稳定性,不损失视觉质量。
- 适合关注动态建模可信度的机器人、强化学习研究者。
视频世界模型虽具备强视觉真实感,但其动态未必真正由动作驱动。本文提出,动作忠实性应基于动作的组合结构来理解,许多具身场景中该结构符合群结构(如导航中的SE(2))。为此,我们形式化了动作条件下的世界建模为状态空间上的群作用,提供超越视觉质量的动态评估准则。为实现此框架,我们提出统一方法,通过合成监督在隐空间施加身份、逆元和组合一致性正则化,无需额外数据采集。此外,引入两个指标:群作用一致性(GAC)和群作用鲁棒性(GAR),用于评估结构正确性与滚动稳定性。大量实验表明,本方法在主流视频世界模型上均显著提升GAC与GAR,且不降低感知质量。
原文摘要 · Abstract (English)
Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness should be understood through the compositional structure of actions, which in many embodied settings follows a group structure (e.g., SE(2) for navigation). Based on this insight, we formalize action-conditioned world modeling as realizing a group action on the state space, providing a principled criterion for evaluating dynamics beyond visual quality. To operationalize this framework, we propose a unified approach that enforces identity, inverse, and composition consistency via latent-space regularization with synthesized supervision, avoiding additional data collection. We further introduce two metrics: Group-Action Consistency (GAC) and Group-Action Robustness (GAR), to evaluate structural correctness and rollout stability. Extensive experimental results show that our method consistently improves both GAC and GAR in state-of-the-art video world models without degrading perceptual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。