arXiv:2508.01766cs.CV2025-08AAAI被引 5

用2D俯视图视觉提示替代语言指令,让智能体更精准导航。

VPN: Visual Prompt Navigation

  • 用户仅需在俯视图上标记路径,无需输入语言指令。
  • 新数据集R2R-VP和R2R-CE-VP提升导航准确率至78.3%。
  • 适合非专业用户,降低导航任务的语义理解门槛。

尽管自然语言常用于引导具身智能体,但其固有的模糊性和冗长性在复杂环境中常阻碍导航效果。为此,我们提出视觉提示导航(VPN),一种新范式:仅通过用户提供的2D俯视图视觉提示来引导智能体导航。该提示聚焦于在场景俯视图中标记视觉导航轨迹,提供直观且空间对齐的指引,无需依赖语言指令,更利于非专业人士使用,减少理解歧义。我们在离散与连续导航设置下构建了VPN任务,并基于现有R2R和R2R-CE数据集扩展出两个新数据集:R2R-VP与R2R-CE-VP。同时引入专用于VPN任务的VPNet基线网络,结合两种数据增强策略——视图级增强(改变初始朝向与提示方向)和轨迹级增强(融合大规模3D场景中的多样化轨迹),以提升导航性能。大量实验评估了视觉提示形式、俯视图地图格式及数据增强策略对导航效果的影响。代码已开源。

原文摘要 · Abstract (English)

While natural language is commonly used to guide embodied agents, the inherent ambiguity and verbosity of language often hinder the effectiveness of language-guided navigation in complex environments. To this end, we propose Visual Prompt Navigation (VPN), a novel paradigm that guides agents to navigate using only user-provided visual prompts within 2D top-view maps. This visual prompt primarily focuses on marking the visual navigation trajectory on a top-down view of a scene, offering intuitive and spatially grounded guidance without relying on language instructions. It is more friendly for non-expert users and reduces interpretive ambiguity. We build VPN tasks in both discrete and continuous navigation settings, constructing two new datasets, R2R-VP and R2R-CE-VP, by extending existing R2R and R2R-CE episodes with corresponding visual prompts. Furthermore, we introduce VPNet, a dedicated baseline network to handle the VPN tasks, with two data augmentation strategies: view-level augmentation (altering initial headings and prompt orientations) and trajectory-level augmentation (incorporating diverse trajectories from large-scale 3D scenes), to enhance navigation performance. Extensive experiments evaluate how visual prompt forms, top-view map formats, and data augmentation strategies affect the performance of visual prompt navigation. The code is available at https://github.com/farlit/VPN.

视觉导航具身智能提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。