arXiv:2604.14785cs.AI2026-04

用镜子测试评估多模态大模型的自我认知能力

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror

论文配图:MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror
图 1 · 摘自论文原文
  • 基于心理学镜像自识实验,设计分层任务评测模型自我认知
  • 顶级模型在基础层面表现仍远低于人类,暴露自我理解缺陷
  • 适合研究通用智能、具身智能与自我意识的学者参考

多模态大语言模型(MLLM)在感知与推理方面取得显著进展,展现出具身智能的潜力。现有基准主要评估模型对外部物体的感知、理解和交互能力,缺乏对自我中心智能的系统性评测。为此,我们提出MirrorBench,一个受心理学经典镜像自识(MSR)测试启发的仿真基准。该基准通过分层递进的任务框架,从基础视觉感知到高层自我表征,评估具身MLLM的自我认知能力。在主流MLLM上的实验表明,即使在最低层级,其表现仍显著落后于人类,揭示了模型在自我指涉理解上的根本局限。本研究连接心理学术语与具身智能,为大型模型中通用智能的涌现提供了原则性评测框架。

原文摘要 · Abstract (English)

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their potential for embodied intelligence. While recent studies have evaluated embodied MLLMs in interactive settings, current benchmarks mainly target capabilities to perceive, understand, and interact with external objects, lacking a systematic evaluation of self-centric intelligence. To address this, we introduce MirrorBench, a simulation-based benchmark inspired by the classical Mirror Self-Recognition (MSR) test in psychology. MirrorBench extends this paradigm to embodied MLLMs through a tiered framework of progressively challenging tasks, assessing agents from basic visual perception to high-level self-representation. Experiments on leading MLLMs show that even at the lowest level, their performance remains substantially inferior to human performance, revealing fundamental limitations in self-referential understanding. Our study bridges psychological paradigms and embodied intelligence, offering a principled framework for evaluating the emergence of general intelligence in large models. Project page: https://fflahm.github.io/mirror-bench-page/.

多模态模型自我认知具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。