提出新方法让大模型在空间认知中理解他人视角,突破传统坐标依赖。
Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks

- 用锚点分解空间推理链,动态融合视听模态信息。
- 在视觉盲区下仍保持42%准确率,优于纯自我中心或全局坐标基线。
- 适合研究具身智能、多智能体协作与跨模态推理的学者参考。
多模态大模型虽在通用推理上表现优异,但其具身空间智能受限于‘笛卡尔幻觉’——过度依赖文本概率分布,缺乏对三维拓扑结构的扎实理解。这一缺陷在多智能体环境中尤为明显,需具备二阶心智理论(ToM):智能体A必须推断智能体B对环境的认知,该认知严格受制于B的物理朝向与感知限制。本文设计了一项新颖的视听任务,要求智能体A预测智能体B对其相对位置的估计。为此,我们提出一个认知性感官瓶颈模块,摒弃固定规则的坐标转换,引入基于锚点的具身空间分解思维链(CoT),引导模型先建立B的局部坐标系,再根据目标是否在视觉视野内动态加权视听信号。大量评估显示,当前多模态大模型在空间对称性和视域外模糊性问题上根本性受限(零样本基准准确率为42%),而我们的感官约束推理链显著优于纯自我中心与全局坐标基线。本工作系统性地评测了感知瓶颈,揭示了多模态大模型空间推理的当前极限,并为具身人工智能中的认知性、模态感知推理奠定了基础范式。
原文摘要 · Abstract (English)
While Multi-Modal Large Language Models (MLLMs) demonstrate impressive capabilities in general reasoning, their embodied spatial intelligence remains hampered by a "Cartesian Illusion" - a reliance on text-based probability distributions that lack grounded, 3D topological understanding. This limitation is starkly exposed in multi-agent environments, which demand more than just scene perception; they require second-order Theory of Mind (ToM). Specifically, an Agent A must be able to infer Agent B's belief about the environment, governed strictly by Agent B's physical orientation and sensory limitations. In this paper, we probe the limits of two-stage spatial inference in MLLMs through a novel audio-visual task: requiring Agent A to predict Agent B's estimation of A's relative location. To solve this, we propose an Epistemic Sensory Bottleneck module that abandons rigid, rule-based coordinate transformations. Instead, we introduce an Anchor-Based Embodied Spatial Decomposition Chain-of-Thought (CoT). This guides the MLLM through a "geometric-to-semantic" projection, forcing it to first establish B's local coordinate system and then dynamically weight visual and auditory modalities based on whether A falls within B's visual frustum. Extensive evaluations reveal that while current MLLMs fundamentally struggle with spatial symmetry and out-of-view ambiguities (establishing a rigorous zero-shot baseline of 42% accuracy), our sensory-bounded reasoning chain robustly outperforms pure egocentric and allocentric baselines. By systematically benchmarking these perceptual bottlenecks, our work exposes the current limits of MLLM spatial reasoning and establishes a foundational paradigm for epistemic, modality-aware inference in Embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。