arXiv:2605.08816cs.AIcs.CY2026-05

测试视觉语言模型能否通过镜子识别自我,揭示其是否具备基于感知与行动的自我认知。

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

论文配图:Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?
图 1 · 摘自论文原文
  • 设计3D镜像任务,让模型从反射中推断身体特征并正确选择目标
  • 强模型能利用镜像信息指导行为,弱模型多误判或无法提取关键信息
  • 揭示模型自指语言不等于真实自我认知,适合研究具身智能的可信性

在动物界,镜像自我识别是高级认知的典型指标,仅出现在部分物种中。我们探究类似功能是否在具身视觉-语言模型(VLM)代理中出现:它们能否在镜中识别自己?为此,我们引入一个受控的3D基准测试,要求第一人称VLM代理根据镜中反射推断隐藏的身体属性,并选择匹配目标,同时避免自我-他人混淆。为区分基于镜像的自我识别与捷径策略,我们测试了移除镜子、误导线索和遮挡反射的情况。还通过镜像探索、时间顺序、自我归因和推理-行动一致性评估决策过程。实验表明,只有较强模型表现出基于镜像的自我识别能力,能利用反射证据指导行动;而较弱模型虽观察镜子,却难以提取自我相关信息或产生误判。语言-视觉冲突进一步显示,仅靠自指语言不足以证明具身自我识别。总体而言,镜像评估可作为诊断工具,判断具身自我定位是否真正根植于感知与行动,而非先验知识、提示迎合或虚构。

原文摘要 · Abstract (English)

In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogous functional capability emerges in embodied vision-language model (VLM) agents: can they recognize themselves in a mirror? We introduce a controlled 3D benchmark where a first-person VLM agent must infer a hidden body attribute from its reflection and select the matching target, while avoiding self-other misattribution. To separate mirror-grounded self-identification from shortcuts, we test mirror removal, misleading cues, and occluded reflections. We also evaluate the decision process through mirror seeking, temporal ordering, self-attribution, and reasoning-action consistency. Our experiments show that mirror-based self-identification emerges mainly in stronger VLMs. These models can use reflected evidence for action, whereas weaker models often inspect the mirror but fail to extract self-relevant information or misattribute their reflection. Language-vision conflict further shows that self-referential language alone is not evidence of grounded self-identification. Overall, mirror-based evaluation provides a diagnostic for whether embodied self-grounding is causally rooted in perception and action rather than priors, prompt compliance, or confabulation.

具身智能自我识别视觉语言模型认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。