arXiv:2605.04327cs.RO2026-05

用语言规则生成逻辑指令,让机器人在户外更安全地自主导航。

From Language to Logic: A Theoretical Architecture for VLM-Grounded Safe Navigation

论文配图:From Language to Logic: A Theoretical Architecture for VLM-Grounded Safe Navigation
图 1 · 摘自论文原文
  • 将自然语言规则转为时序逻辑,指导机器人实时规划与避障。
  • 通过环境感知动态调整路径,满足多类安全约束和偏好要求。
  • 适合需高安全性的户外机器人应用,如无人巡检或救援任务。

我们提出一种架构,将人类提供的高层次安全规则与操作者对语义的偏好融入非结构化户外环境中的自主机器人导航。在该方法中,自然语言规则被转化为信号时序逻辑(STL)规范,用于运行时的规划与导航。持久的、以环境为中心的规则和地形偏好被嵌入二维代价地图,而具有时间动态性的要求则以STL规范形式表达,并在运行时进行监控。我们假设使用视觉-语言模型(VLMs)实现零样本场景理解,建立人类指令、语义特征与环境约束之间的映射。在此框架下,我们构建了一个可满足一组编码于STL中的规范及软性操作者偏好的导航模型,通过嵌入环境属性的正式满足度量与运行时监控实现目标。

原文摘要 · Abstract (English)

We propose an architecture for integrating high-level, human-provided safety rules and operator-aligned semantic preferences into autonomous robot navigation in unstructured outdoor environments. In our approach, natural-language rules are translated into Signal Temporal Logic (STL) specifications that guide planning and navigation during runtime. Persistent, environment-centric rules and terrain preferences are grounded into a 2D cost map, while temporally dynamic requirements are expressed as STL specifications to be monitored during runtime. We hypothesize the use of Vision-Language Models (VLMs) for zero-shot scene understanding, enabling mapping between human instructions, semantic features, and environmental constraints. Within this framework, we construct an illustrative navigation model that is designed to satisfy a set of STL-encoded specifications and soft operator preferences through formal satisfaction metrics embedded into environmental properties and runtime monitoring.

机器人导航语言逻辑安全控制VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。