arXiv:2512.06232cs.CVcs.LG2025-12

用儿童视角视频训练模型,仍难学会物理直觉

Opinion: Learning Intuitive Physics May Require More than Visual Data

  • 在儿童日常视觉数据上预训练,模拟发展过程
  • 仅用0.01%的视频量,性能未显著提升
  • 当前模型需更复杂数据或架构支持物理直觉

人类通过构建基于物理直觉的内部模型来高效应对世界。尽管当前深度学习模型在大量互联网视频数据上训练,但在直观物理任务上的表现仍远低于人类水平。本文探讨数据分布而非数据量是否是关键。我们在SAYCam——一个部分反映三个儿童日常视觉体验的、发展性真实的视角视频数据集上,预训练了视频联合嵌入预测架构(V-JEPA)模型。结果显示,该数据量仅为顶尖模型所用数据的0.01%,在IntPhys2基准测试上并未带来显著性能提升。研究表明,仅使用发展性真实数据不足以使现有模型学习到支持直观物理的认知表征。因此,单纯改变视觉数据的数量和分布可能不足以构建具备人工物理直觉的系统。

原文摘要 · Abstract (English)

Humans expertly navigate the world by building rich internal models founded on an intuitive understanding of physics. Meanwhile, despite training on vast quantities of internet video data, state-of-the-art deep learning models still fall short of human-level performance on intuitive physics benchmarks. This work investigates whether data distribution, rather than volume, is the key to learning these principles. We pretrain a Video Joint Embedding Predictive Architecture (V-JEPA) model on SAYCam, a developmentally realistic, egocentric video dataset partially capturing three children's everyday visual experiences. We find that training on this dataset, which represents 0.01% of the data volume used to train SOTA models, does not lead to significant performance improvements on the IntPhys2 benchmark. Our results suggest that merely training on a developmentally realistic dataset is insufficient for current architectures to learn representations that support intuitive physics. We conclude that varying visual data volume and distribution alone may not be sufficient for building systems with artificial intuitive physics.

物理直觉视频建模发展认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。