提出新基准与模型,让AI像人一样在脑中规划路线。
Mind over Space: Can Multimodal Large Language Models Mentally Navigate?
- 用长视频构建分层认知地图,逐步生成基于地标路径计划。
- 主流大模型零样本下空间规划准确率随距离急剧下降。
- 新模型NavMind通过细粒度认知地图实现更优导航能力,适合智能体研发者。
尽管多模态大模型在具身智能体中广泛应用,其能力仍局限于从即时观察进行反应式规划,难以应对长时空尺度的空间推理。认知科学表明,生物智能依赖于“心理导航”:基于经验构建空间表征,并在行动前进行路径心理模拟。为弥合人工智能与生物智能的差距,我们提出Video2Mental,首个评估多模态大模型心理导航能力的基准。任务要求从长时自我中心视频构建分层认知地图,并逐步生成基于地标路径计划,规划准确性通过仿真器物理交互验证。基准测试显示,心理导航能力不会自然涌现于标准预训练;前沿多模态大模型在零样本下对结构化空间表征表现极差,且规划准确率随远距离迅速衰减。为此,我们提出NavMind,一种将心理导航内化的推理模型,以可学习的细粒度认知地图作为中间表示。通过难度分层的渐进式监督微调策略,NavMind有效连接原始感知与结构化规划。实验表明,NavMind显著优于前沿商业及空间感知多模态大模型。
原文摘要 · Abstract (English)
Despite the widespread adoption of MLLMs in embodied agents, their capabilities remain largely confined to reactive planning from immediate observations, consistently failing in spatial reasoning across extensive spatiotemporal scales. Cognitive science reveals that Biological Intelligence (BI) thrives on "mental navigation": the strategic construction of spatial representations from experience and the subsequent mental simulation of paths prior to action. To bridge the gap between AI and BI, we introduce Video2Mental, a pioneering benchmark for evaluating the mental navigation capabilities of MLLMs. The task requires constructing hierarchical cognitive maps from long egocentric videos and generating landmark-based path plans step by step, with planning accuracy verified through simulator-based physical interaction. Our benchmarking results reveal that mental navigation capability does not naturally emerge from standard pre-training. Frontier MLLMs struggle profoundly with zero-shot structured spatial representation, and their planning accuracy decays precipitously over extended horizons. To overcome this, we propose \textbf{NavMind}, a reasoning model that internalizes mental navigation using explicit, fine-grained cognitive maps as learnable intermediate representations. Through a difficulty-stratified progressive supervised fine-tuning paradigm, NavMind effectively bridges the gap between raw perception and structured planning. Experiments demonstrate that NavMind achieves superior mental navigation capabilities, significantly outperforming frontier commercial and spatial MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。