arXiv:2605.15466cs.CV2026-05

让模型关注物体互动,提升因果推理能力。

Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction

论文配图:Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction
图 1 · 摘自论文原文
  • 用运动中心的掩码策略聚焦物理交互事件。
  • 在CLEVRER上达14.26%因果推理准确率,远超基线3.22%。
  • 适合做物理世界建模、因果推理的研究者使用。

从无标签视频中学习预测性世界模型是人工智能的基础挑战。尽管联合嵌入预测架构(JEPA)在语义分类上取得新基准,但常缺乏物理感知,无法捕捉下游推理所需的因果动态。我们假设这源于标准的基于块的掩码策略,其优先考虑视觉纹理而非罕见但信息量大的运动事件。为此提出交互感知JEPA(IA-JEPA),采用自监督运动中心掩码策略,聚焦于发生碰撞或动量传递的物体。通过专门针对这些互动实体,强制模型重建潜在轨迹而非静态背景特征。在CLEVRER基准测试中,IA-JEPA在因果推理任务上达到14.26%准确率,显著优于标准块掩码基线的3.22%。关键的是,我们证明了IA-JEPA打破了标准自监督的“静态偏差”,使潜在空间熵值提升10%,并线性化物理能量(R²=0.43)。该交互偏见还推广至真实人类动作(Something-Something V2)和零样本物理谜题(PHYRE-Lite)。结果表明,这为构建开始内化物理世界因果结构的可扩展自监督世界模型提供了可行路径。

原文摘要 · Abstract (English)

Learning predictive world models from unlabelled video is a foundational challenge in artificial intelligence. While Joint Embedding Predictive Architectures (JEPA) have set new benchmarks in semantic classification, they often remain physics-blind, failing to capture the causal dynamics necessary for downstream reasoning. We hypothesize that this stems from standard patch-based masking strategies, which prioritize visual texture over rare but informative kinematic events. We propose Interaction-Aware JEPA (IA-JEPA), which utilizes a self-supervised motion-centric masking strategy to prioritize physical interactions. By specifically targeting entities engaged in collisions or momentum transfers, we force the architecture to reconstruct latent trajectories rather than static background features. Evaluated on the CLEVRER benchmark, IA-JEPA achieves 14.26% accuracy on causal reasoning tasks, a significant lead over the 3.22% achieved by standard patch-masked baselines. Crucially, we demonstrate that IA-JEPA breaks the "static bias" of standard self-supervision by inducing a higher-entropy, more discriminative latent space (+10% entropy gain) that linearizes physical energy ($R^2=0.43$). We show that this interaction bias generalizes to real-world human actions (Something-Something V2) and zero-shot physical puzzles (PHYRE-Lite). Our results provide a scalable, fully self-supervised path toward building foundational world models that begin to internalize the causal structure of the physical world.

世界模型因果推理视频预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。