arXiv:2508.01715cs.RO2025-08中稿 · and presented at t…

用视觉语言模型实现零样本地形可通行性估计,探索其可行性和挑战

Towards Zero-Shot Terrain Traversability Estimation: Challenges and Opportunities

  • 基于人类标注数据构建小规模可通行性评估集
  • 当前基础模型在零样本下表现不稳定,但具研究价值
  • 适合对自主导航和多模态感知感兴趣的开发者

地形可通行性估计对自主机器人至关重要,尤其在非结构化环境中,视觉线索与推理能力尤为关键。尽管视觉语言模型(VLMs)为零样本估计提供了可能,但该问题本质上仍属不适定。为此,我们引入一个由人工标注的水体可通行性评分小数据集,结果显示尽管判断存在主观性,但人类评估者间仍有一定共识。此外,我们提出一种简单流水线,整合VLM实现零样本可通行性估计。实验表明结果参差不齐,暗示当前基础模型尚不适合实际部署,但为后续研究提供了宝贵洞见。

原文摘要 · Abstract (English)

Terrain traversability estimation is crucial for autonomous robots, especially in unstructured environments where visual cues and reasoning play a key role. While vision-language models (VLMs) offer potential for zero-shot estimation, the problem remains inherently ill-posed. To explore this, we introduce a small dataset of human-annotated water traversability ratings, revealing that while estimations are subjective, human raters still show some consensus. Additionally, we propose a simple pipeline that integrates VLMs for zero-shot traversability estimation. Our experiments reveal mixed results, suggesting that current foundation models are not yet suitable for practical deployment but provide valuable insights for further research.

机器人视觉语言模型地形估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。