用视觉问答评测大模型是否具备身体认知能力
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
- 设计双任务基准,从第一视角交互中建模世界变化
- 模型在长时序任务中表现远低于人类,差距随交互长度增大
- 暴露模型对人视角的依赖和右利手偏好,适合研究具身智能
具身认知认为智能源于感官运动交互而非被动观察。我们提出ENACT基准,将具身认知评估转化为视觉问答中的第一视角交互世界建模。该基准基于部分可观测马尔可夫决策过程(POMDP),以场景图变化为动作,包含正向世界建模(根据动作重排序观测)与逆向世界建模(根据观测重排序动作)两个互补任务。虽概念简单,但求解需具备功能识别、动作效应推理、具身意识及长时交互记忆等能力,且避免低层图像生成干扰。我们构建可扩展的流水线,从机器人仿真(BEHAVIOR)生成8,972组问答对,涵盖长时序家庭活动。实验显示,前沿视觉语言模型与人类间存在性能差距,且差距随交互时长增加;模型在逆向任务中表现优于正向任务,并表现出人本偏见,如偏好右手动作,以及在相机内参或视角偏离人类视觉时性能下降。
原文摘要 · Abstract (English)
Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodied manner, exhibit signs of embodied cognition? We introduce ENACT, a benchmark that casts evaluation of embodied cognition as world modeling from egocentric interaction in a visual question answering (VQA) format. Framed as a partially observable Markov decision process (POMDP) whose actions are scene graph changes, ENACT comprises two complementary sequence reordering tasks: forward world modeling (reorder shuffled observations given actions) and inverse world modeling (reorder shuffled actions given observations). While conceptually simple, solving these tasks implicitly demands capabilities central to embodied cognition-affordance recognition, action-effect reasoning, embodied awareness, and interactive, long-horizon memory from partially observable egocentric input, while avoiding low-level image synthesis that could confound the evaluation. We provide a scalable pipeline that synthesizes QA pairs from robotics simulation (BEHAVIOR) and evaluates models on 8,972 QA pairs spanning long-horizon home-scale activities. Experiments reveal a performance gap between frontier VLMs and humans that widens with interaction horizon. Models consistently perform better on the inverse task than the forward one and exhibit anthropocentric biases, including a preference for right-handed actions and degradation when camera intrinsics or viewpoints deviate from human vision. Website at https://enact-embodied-cognition.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。