提出统一认知框架,整合机器世界模型的七大认知能力
Human Cognition in Machines: A Unified Perspective of World Models

- 构建包含记忆、感知、语言等七项认知功能的统一框架
- 发现动机与元认知研究严重不足,需加强内在驱动力建模
- 引入知识型世界模型,推动科学发现类智能体发展
本报告从认知功能角度区分世界模型的研究工作。许多研究宣称其模型具备接近人类的认知能力。评估这些声称需要基于人类与机器认知理论的基本原则。在迈向类人世界模型的过程中,我们提出一个全面整合记忆、感知、语言、推理、想象、动机和元认知七项认知功能的统一框架,并识别现有研究中的空白,为未来前沿提供指引。特别地,我们发现动机(尤其是内在动机)和元认知仍严重缺乏研究,据此提出结合主动推断与全局工作空间理论的具体改进方向。此外,我们引入了知识型世界模型这一新类别,涵盖以结构化知识为基础进行科学发现的智能体框架。该分类体系应用于视频、具身及知识型世界模型,揭示了以往分类体系未覆盖的研究方向。
原文摘要 · Abstract (English)
This report of world models distinguishes prior works by the cognitive functions they innovate. Many works claim an almost human-like cognitive capability in their world models. To evaluate these claims requires a proper grounding in first principles from human and machine cognition theory. In moving towards human-like world models we present a conceptual unified framework for world models that fully incorporates all the cognitive functions (i.e., memory, perception, language, reasoning, imagining, motivation, and metacognition) and identify gaps in existing research as a guide for future states of the art. In particular, we find that motivation (especially intrinsic motivation) and metacognition remain drastically under-researched, and we propose concrete directions to address these gaps informed by active inference and global workspace theory. We also introduce epistemic world models, a new category encompassing agent frameworks for scientific discovery that operate over structured knowledge. Our taxonomy, applied to video, embodied, and epistemic world models, suggests research directions where prior taxonomies have not.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。