让AI导航不仅知路,更懂规则,避免闯红灯、逆行等违规行为。
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification

- 引入语义推理与几何修正模块,让AI理解交通规则而非仅看路径。
- 在29,000节点城市环境中测试,177类规则约束下提升导航合规性。
- 零样本通用模块可适配现有模型,适合智能驾驶与服务机器人应用。
随着具身智能向真实世界部署演进,视觉语言导航(VLN)任务的成功标准正从单纯可达性转向社会合规性。然而,当前智能体陷入“目标驱动陷阱”,过度关注物理可达性(“能否走?”)而忽视语义规则(“是否允许走?”),常忽略细微的规制约束。为此,我们构建了首个大规模城市级合规导航基准Rule-VLN,涵盖29,000个节点的环境,在8,000个受限节点中注入177种多样化的规则类别,分四个课程层级设置挑战。同时提出语义导航修正模块(SNRM),一种通用的零样本模块,用于赋予预训练智能体安全意识。SNRM结合粗到细的视觉感知VLM框架与认知型心理地图,实现动态绕行规划。实验表明,尽管Rule-VLN对最先进模型构成挑战,但SNRM显著恢复导航能力,将合规率下降(CVR)降低19.26%,同时提升成功率(TC)5.97%。
原文摘要 · Abstract (English)
As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachability to social compliance. However, current agents suffer from a "goal-driven trap", prioritizing physical geometry ("can I go?") over semantic rules ("may I go?"), frequently overlooking subtle regulatory constraints. To bridge this gap, we establish Rule-VLN, the first large-scale urban benchmark for rule-compliant navigation. Spanning a massive 29k-node environment, it injects 177 diverse regulatory categories into 8k constrained nodes across four curriculum levels, challenging agents with fine-grained visual and behavioral constraints. We further propose the Semantic Navigation Rectification Module (SNRM), a universal, zero-shot module designed to equip pre-trained agents with safety awareness. SNRM integrates a coarse-to-fine visual perception VLM framework with an epistemic mental map for dynamic detour planning. Experiments demonstrate that while Rule-VLN challenges state-of-the-art models, SNRM significantly restores navigation capabilities, reducing CVR by 19.26% and boosting TC by 5.97%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。