测试视觉语言模型在安全场景下的空间认知能力,发现其决策依赖文本偏见而非真实环境。
Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

- 构建探索-地图-记忆-决策框架,量化模型的空间理解能力
- 模型在黑暗中推理能力下降,但对颜色纹理篡改不敏感
- 记忆机制与人类认知差异大,存在不可预测的安全风险
本文提出理论空间框架(ToS)的扩展版本——面向安全关键场景的探索-地图-记忆-决策(EMRD)流水线,评估好奇心驱动的视觉语言模型(VLMs)在部分可观测环境下的空间理解能力。通过环境覆盖度与时间效率衡量探索能力(Explore),以空间保真度评估地图构建(Map),采用心理测量指标评估记忆持久性(Remember),并利用焦点点指标测量认知决策能力(Decide)。实验表明,模型常依据预训练文本先验选择疏散点,缺乏空间依据;在低光条件下空间推理性能显著下降,但不受纹理与色彩篡改影响。结果揭示VLM记忆机制与人类认知本质不同,可能引发严重误判风险。
原文摘要 · Abstract (English)
Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI techniques are increasingly applied to safety-critical scenarios, it is crucial to understand whether VLMs possess robust spatial memory and make reliable decisions. In this paper, we assess whether VLMs' decisions are based on physical evidence or are corrupted by visual-language biases, if their memory processes align with human cognitive patterns, and how they respond to environmental hazards. We extend the ToS framework into a safety-critical, goal-driven pipeline, named Explore, Map, Remember, and Decide (EMRD). We then quantify Exploration Competence (Explore) through metrics of environmental coverage and temporal efficiency, assess Spatial Fidelity (Map), evaluate, with a suite of psychological metrics, Memory Persistence (Remember), and measure, using focal-point metrics, Cognitive Decision-Making (Decide). Our results show that in terms of decision-making capabilities, VLMs frequently select evacuation points based on pre-trained textual priors while lacking the spatial grounding to justify their choices. We also show that spatial reasoning degrades in low-light conditions, but it is not affected by texture and colour tampering. Our findings suggest that VLM memory fundamentally diverges from human cognition, creating unpredictable risks of misalignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。