arXiv:2604.10333cs.AIcs.CV2026-04被引 3

用零样本世界模型模拟儿童高效学习物理世界的能力

Zero-shot World Models Are Developmentally Efficient Learners

论文配图:Zero-shot World Models Are Developmentally Efficient Learners
图 1 · 摘自论文原文
  • 基于稀疏时序分解预测器,分离视觉与动态信息
  • 仅凭一个孩子的一次经验,快速掌握多个物理理解任务
  • 模拟儿童发展行为模式,适合研究认知科学与高效AI

幼儿展现出对物理世界的早期理解能力,如深度估计、运动感知、物体连贯性判断及交互识别等。他们虽训练数据极少,却具备极高的数据效率与泛化能力,这对当前最先进的人工智能系统仍是巨大挑战。本文提出一种新计算假说——零样本视觉世界模型(ZWM),其基于三个原则:稀疏的时序因子化预测器,将外观与动态解耦;通过近似因果推断实现零样本估计;通过推理组合构建更复杂能力。我们证明,仅需一个孩子的一次第一人称体验数据,ZWM即可迅速在多个物理理解基准上生成胜任能力。该模型还广泛复现了儿童发展的行为特征,并生成类脑内部表征。本工作为从人类规模数据中实现高效灵活学习提供了蓝图,既推进了对儿童早期物理理解的计算解释,也为数据高效人工智能系统指明了路径。

原文摘要 · Abstract (English)

Young children demonstrate early abilities to understand their physical world, estimating depth, motion, object coherence, interactions, and many other aspects of physical scene understanding. Children are both data-efficient and flexible cognitive systems, creating competence despite extremely limited training data, while generalizing to myriad untrained tasks -- a major challenge even for today's best AI systems. Here we introduce a novel computational hypothesis for these abilities, the Zero-shot Visual World Model (ZWM). ZWM is based on three principles: a sparse temporally-factored predictor that decouples appearance from dynamics; zero-shot estimation through approximate causal inference; and composition of inferences to build more complex abilities. We show that ZWM can be learned from the first-person experience of a single child, rapidly generating competence across multiple physical understanding benchmarks. It also broadly recapitulates behavioral signatures of child development and builds brain-like internal representations. Our work presents a blueprint for efficient and flexible learning from human-scale data, advancing both a computational account for children's early physical understanding and a path toward data-efficient AI systems.

世界模型认知科学零样本学习数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。