arXiv:2511.17384cs.ROcs.CV2025-11被引 3

首个动态工业场景导航基准,测试智能体空间推理与安全行为

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

  • 构建12个高保真动态仓库场景,评估智能体主动规划能力
  • 封闭模型表现领先,但普遍存在路径规划差、避障弱等问题
  • 适合关注具身智能、安全导航与动态环境建模的研究者

尽管视觉大语言模型(VLLMs)在具身智能领域展现出巨大潜力,但在空间推理方面仍面临严峻挑战。现有具身评测基准多聚焦于静态、被动的家居环境,仅评估单一能力,难以反映特定领域中交互性与动态复杂性的综合表现。为此,我们提出IndustryNav,首个面向动态工业导航的主动空间推理评测基准。该基准包含12个手动构建的高保真Unity仓库场景,涵盖动态物体与人类移动。我们设计了一种零样本点目标导航流程,有效融合视点视觉与全局里程计,评估局部-全局协同规划能力。同时引入“碰撞率”与“预警率”指标,量化安全行为。对十四种先进VLLMs(如GPT-5.2、Claude-4.6、Gemini-3)的全面评估表明,封闭源模型保持稳定优势;然而所有智能体在鲁棒路径规划、碰撞规避与主动探索方面均存在明显不足。这凸显了具身研究亟需从被动感知转向要求稳定规划、主动探索与安全行为的动态真实场景。

原文摘要 · Abstract (English)

While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reasoning. Existing embodied benchmarks largely focus on passive, static household environments and evaluate isolated capabilities, failing to capture holistic performance in interactive and dynamic complexity of specific domains. To fill this gap, we present IndustryNav, the first dynamic industrial navigation benchmark for active spatial reasoning. IndustryNav leverages 12 manually created, high-fidelity Unity warehouse scenarios featuring dynamic objects and human movement. We proposes a zero-shot PointGoal navigation pipeline that effectively combines egocentric vision with global odometry to assess holistic local-global planning. Furthermore, we introduce the "collision rate" and "warning rate" metrics to measure safety-oriented behaviors. A comprehensive study of fourteen state-of-the-art VLLMs (including models such as GPT-5.2, Claude-4.6, and Gemini-3) reveals that closed-source models maintain a consistent advantage; however, all agents exhibit notable deficiencies in robust path planning, collision avoidance and active exploration. This highlights a critical need for embodied research to move beyond passive perception and toward tasks that demand stable planning, active exploration, and safe behavior in vivid, dynamic environments.

具身智能空间推理动态导航安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。