arXiv:2503.21668cs.AIcs.CV2025-03被引 4

从认知科学出发,评估AI对物体理解的核心能力。

Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI

  • 融合格式塔、具身认知等理论,梳理物体理解的关键能力。
  • 现有基准仅能检测局部物体能力,无法评估功能整合。
  • 提出新评测方法,推动AI向真实世界物体理解迈进。

世界模型的核心成分之一是‘直觉物理’——对物体、空间和因果关系的理解。这一能力使我们能够预测事件、规划行动并导航环境,均依赖于对物体整体性的综合感知。尽管其重要性突出,目前尚无统一的物体概念框架,但多个理论体系提供了有益见解。本文第一部分综述了物体理解研究中的主要理论框架:格式塔心理学、具身认知和发育心理学,明确了各框架赋予物体理解的核心能力及其在生物体世界模型构建中的功能角色。鉴于物体理解在世界建模中的基础地位,理解它对人工智能同样至关重要。第二部分评估当前人工智能范式在物体理解上的实现与测试方式,相较于认知科学的进展。我们将人工智能范式定义为对象理解的构念、研究方法、数据使用及评估技术的结合。研究发现,虽然现有基准可识别出人工智能系统对物体理解的某些孤立方面有所建模,却无法检测其在这些能力之间缺乏功能整合的问题,因而未能真正解决物体理解挑战。最后,本文探索了与本文提出的整合性物体理解观相一致的新评估方法,这些方法有望推动人工智能从孤立的物体能力迈向具备真实世界情境下真正物体理解的通用智能。

原文摘要 · Abstract (English)

One of the core components of our world models is 'intuitive physics' - an understanding of objects, space, and causality. This capability enables us to predict events, plan action and navigate environments, all of which rely on a composite sense of objecthood. Despite its importance, there is no single, unified account of objecthood, though multiple theoretical frameworks provide insights. In the first part of this paper, we present a comprehensive overview of the main theoretical frameworks in objecthood research - Gestalt psychology, enactive cognition, and developmental psychology - and identify the core capabilities each framework attributes to object understanding, as well as what functional roles they play in shaping world models in biological agents. Given the foundational role of objecthood in world modelling, understanding objecthood is also essential in AI. In the second part of the paper, we evaluate how current AI paradigms approach and test objecthood capabilities compared to those in cognitive science. We define an AI paradigm as a combination of how objecthood is conceptualised, the methods used for studying objecthood, the data utilised, and the evaluation techniques. We find that, whilst benchmarks can detect that AI systems model isolated aspects of objecthood, the benchmarks cannot detect when AI systems lack functional integration across these capabilities, not solving the objecthood challenge fully. Finally, we explore novel evaluation approaches that align with the integrated vision of objecthood outlined in this paper. These methods are promising candidates for advancing from isolated object capabilities toward general-purpose AI with genuine object understanding in real-world contexts.

物体理解认知科学评估方法通用AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。