arXiv:2603.22800cs.RO2026-03被引 1

用视觉语言模型实现零样本机器人导航,自动评估地形风险并减少计算量。

CATNAV: Cached Vision-Language Traversability for Efficient Zero-Shot Robot Navigation

  • 通过视觉语义缓存复用相似场景的风险评估,降低85.7%的在线查询。
  • 在五项任务中达成10%更高成功率,行为违规减少33%。
  • 适合需要快速适应新环境的机器人系统,尤其适用于四足机器人。

导航非结构化环境需根据机器人的物理能力评估通行风险,这一挑战因机体差异而异。我们提出CATNAV,一种基于多模态大模型的零样本、体感感知成本图生成框架,无需任务特定训练即可完成风险评估。引入视觉语义缓存机制,检测场景新颖性并复用先前相似帧的风险评估,使在线视觉语言模型(VLM)查询减少85.7%。此外,设计基于VLM的轨迹选择模块,通过视觉推理评估路径提案,在满足行为约束条件下选择最安全路径。我们在四足机器人上于室内外非结构化环境中评估CATNAV,对比现有最先进的视觉-语言-动作基线。在五个导航任务中,CATNAV平均到达率提升10个百分点,行为约束违反次数减少33%。

原文摘要 · Abstract (English)

Navigating unstructured environments requires assessing traversal risk relative to a robot's physical capabilities, a challenge that varies across embodiments. We present CATNAV, a cost-aware traversability navigation framework that leverages multimodal LLMs for zero-shot, embodiment-aware costmap generation without task-specific training. We introduce a visuosemantic caching mechanism that detects scene novelty and reuses prior risk assessments for semantically similar frames, reducing online VLM queries by 85.7%. Furthermore, we introduce a VLM-based trajectory selection module that evaluates proposals through visual reasoning to choose the safest path given behavioral constraints. We evaluate CATNAV on a quadruped robot across indoor and outdoor unstructured environments, comparing against state-of-the-art vision-language-action baselines. Across five navigation tasks, CATNAV achieves 10 percentage point higher average goal-reaching rate and 33% fewer behavioral constraint violations.

机器人导航多模态模型零样本成本图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。