arXiv:2601.07375cs.CL2026-01ACL被引 1

用地图数据评估导航指令效果,无需视觉信息和训练。

GROKE: Vision-Free Navigation Instruction Evaluation via Graph Reasoning on OpenStreetMap

  • 基于地图知识的分层大模型框架,通过结构化空间信息判断指令可行性。
  • 在Map2Seq数据集上导航误差降低68.5%,优于启发式与采样基线。
  • 适合关注导航指令语义有效性、追求可解释性评估的研究者。

导航指令评估在视觉-语言导航(VLN)研究中仍是难题。传统基于参考的指标如BLEU和ROUGE无法捕捉空间指令的功能效用,即是否真正引导导航者到达目标地点。现有VLN智能体虽可作为评估器,但依赖高保真视觉模拟器带来许可与计算成本,感知误差也干扰语言质量判断。本文提出GROKE(基于开放街图知识的图推理导航指令评估),一种无需训练、不依赖视觉的分层大模型框架,利用开放街图(OSM)数据评估导航指令。系统消融实验表明,结构化JSON与文本格式的空间信息表达显著优于网格与视觉图表示。其分层架构结合子指令规划与拓扑图导航,在Map2Seq数据集上相较启发式与采样基线将导航误差降低68.5%。智能体执行成功率、轨迹保真度与决策模式成为基于可见地标与拓扑结构的功能可导航性代理指标,构建了无视觉依赖、可扩展且可解释的评估范式。代码与数据已公开于https://anonymous.4open.science/r/groke。

原文摘要 · Abstract (English)

The evaluation of navigation instructions remains a persistent challenge in Vision-and-Language Navigation (VLN) research. Traditional reference-based metrics such as BLEU and ROUGE fail to capture the functional utility of spatial directives, specifically whether an instruction successfully guides a navigator to the intended destination. Although existing VLN agents could serve as evaluators, their reliance on high-fidelity visual simulators introduces licensing constraints and computational costs, and perception errors further confound linguistic quality assessment. This paper introduces GROKE(Graph-based Reasoning over OSM Knowledge for instruction Evaluation), a vision-free training-free hierarchical LLM-based framework for evaluating navigation instructions using OpenStreetMap data. Through systematic ablation studies, we demonstrate that structured JSON and textual formats for spatial information substantially outperform grid-based and visual graph representations. Our hierarchical architecture combines sub-instruction planning with topological graph navigation, reducing navigation error by 68.5% compared to heuristic and sampling baselines on the Map2Seq dataset. The agent's execution success, trajectory fidelity, and decision patterns serve as proxy metrics for functional navigability given OSM-visible landmarks and topology, establishing a scalable and interpretable evaluation paradigm without visual dependencies. Code and data are available at https://anonymous.4open.science/r/groke.

导航评估大模型地图推理无视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。