arXiv:2605.10376cs.CV2026-05

测试视觉语言模型在3D场景中执行指令的导航能力,发现其在复杂任务下表现显著下降。

SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation

论文配图:SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation
图 1 · 摘自论文原文
  • 构建三层次基准,评估模型在局部交互中的空间推理能力
  • 2472个场景中,多步指令下性能下降超40%
  • 适合研究具身智能、视觉语言导航与动作生成的学者

视觉语言模型(VLMs)在多模态感知和语言理解方面进展迅速,但其是否能在3D数字环境中可靠地将语言映射为空间连贯、可执行的动作仍不明确。我们提出SleepWalk,一个用于评估单场景3D世界中指令引导轨迹预测的基准。该环境由文本描述生成并筛选出可导航区域。与以往关注跨房间长程探索的基准不同,SleepWalk聚焦于局部化、以交互为中心的具身推理:给定渲染的视觉观测和自然语言指令,模型需预测符合场景几何、避免碰撞且终止于可执行位置的轨迹。基准涵盖多样室内与室外环境,并按空间和时间复杂度分为三个层级,支持对语义接地在组合复杂度增加下的细粒度分析。采用标准化点判评估协议,在2,472个精心筛选的3D环境中对三种前沿VLM进行评估,每个场景有九条指令。结果揭示系统性失败,尤其在遮挡、交互约束和多步指令下:任务难度越高,性能越差。总体而言,当前VLM可部分生成空间连贯、合理可执行且与意图一致的轨迹。通过在可控而可扩展的设置中暴露缺陷,SleepWalk为推进具身多模态推理、具身规划、视觉语言导航及3D环境中具备动作能力的智能体提供了关键基准。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliably ground language into spatially coherent, plausibly executable actions in 3D digital environments. We introduce SleepWalk, a benchmark for evaluating instruction-grounded trajectory prediction in single-scene 3D worlds generated from textual scene descriptions and filtered for navigability. Unlike prior navigation benchmarks centered on long-range exploration across rooms, SleepWalk targets localized, interaction-centric embodied reasoning: given rendered visual observations and a natural-language instruction, a model must predict a trajectory that respects scene geometry, avoids collisions, and terminates at an action-compatible location. The benchmark covers diverse indoor and outdoor environments and organizes tasks into three tiers of spatial and temporal difficulty, enabling fine-grained analysis of grounding under increasing compositional complexity. Using a standardized pointwise judge-based evaluation protocol, we evaluate three frontier VLMs on 2,472 curated 3D environments with nine instructions per scene. Results reveal systematic failures in grounded spatial reasoning, especially under occlusion, interaction constraints, and multi-step instructions: performance drops as the difficulty level of the tasks increase. In general, current VLMs can somewhat produce trajectories that are simultaneously spatially coherent, plausibly executable, and aligned with intended actions. By exposing failures in a controlled yet scalable setting, SleepWalk provides a critical benchmark for advancing grounded multimodal reasoning, embodied planning, vision-language navigation, and action-capable agents in 3D environments.

视觉语言导航具身智能3D推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。