arXiv:2502.11831cs.CVcs.AI2025-02被引 72

模型通过自监督学习自然视频,自发掌握物体恒常性等物理直觉。

Intuitive physics understanding emerges from self-supervised pretraining on natural videos

  • 在抽象表示空间中预测视频缺失区域,可习得物理直觉。
  • 仅用一周真实视频训练即达到高于随机水平的性能。
  • 无需预设知识模块,也能理解物体持续存在与形状不变。

我们研究了通用深度神经网络模型在自然视频上进行自监督预训练后,对直观物理理解的涌现现象。通过违反预期框架评估,发现基于学习表示空间预测视频结果的模型能够理解物体恒常性、形状一致性等直观物理属性。相比之下,在像素空间进行视频预测或使用多模态大语言模型(通过文本推理)的表现接近随机水平。这些对比表明,同时学习抽象表示空间并预测感官输入缺失部分(类似预测编码机制),足以获得直观物理理解;即使仅用一周独特视频训练,模型也表现优于随机。这挑战了核心知识——即先天认知系统——需预先设定才能形成直观物理理解的观点。

原文摘要 · Abstract (English)

We investigate the emergence of intuitive physics understanding in general-purpose deep neural network models trained to predict masked regions in natural videos. Leveraging the violation-of-expectation framework, we find that video prediction models trained to predict outcomes in a learned representation space demonstrate an understanding of various intuitive physics properties, such as object permanence and shape consistency. In contrast, video prediction in pixel space and multimodal large language models, which reason through text, achieve performance closer to chance. Our comparisons of these architectures reveal that jointly learning an abstract representation space while predicting missing parts of sensory input, akin to predictive coding, is sufficient to acquire an understanding of intuitive physics, and that even models trained on one week of unique video achieve above chance performance. This challenges the idea that core knowledge -- a set of innate systems to help understand the world -- needs to be hardwired to develop an understanding of intuitive physics.

自监督学习直观物理视频预测认知涌现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。