用检索增强生成用户真实需求指令,提升机器人导航数据质量。
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM
- 构建场景分层描述树,从全局到局部理解3D环境
- 模拟多种用户角色生成200万条多样化导航指令
- 适合需要真实用户表达风格的智能体导航研究
视觉-语言导航(VLN)是智能体在三维环境中根据自然语言指令移动的关键能力。高性能导航模型依赖大量训练数据,但人工标注成本高昂,严重制约该领域发展。此前方法通过轨迹视频生成步骤化指令,但与用户简洁描述目的地或表达具体需求的沟通习惯不符,且忽略全局上下文和高层任务规划。为此,我们提出NavRAG——一种基于检索增强生成(RAG)的框架,用于生成符合用户真实需求的导航指令。该框架利用大语言模型(LLM)构建从全局布局到局部细节的层次化场景描述树,模拟不同用户角色及特定需求,从场景树中检索并生成多样化指令。我们在861个场景中人工标注了超过200万条导航指令,并评估了数据质量与模型导航性能。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) is an essential skill for embodied agents, allowing them to navigate in 3D environments following natural language instructions. High-performance navigation models require a large amount of training data, the high cost of manually annotating data has seriously hindered this field. Therefore, some previous methods translate trajectory videos into step-by-step instructions for expanding data, but such instructions do not match well with users' communication styles that briefly describe destinations or state specific needs. Moreover, local navigation trajectories overlook global context and high-level task planning. To address these issues, we propose NavRAG, a retrieval-augmented generation (RAG) framework that generates user demand instructions for VLN. NavRAG leverages LLM to build a hierarchical scene description tree for 3D scene understanding from global layout to local details, then simulates various user roles with specific demands to retrieve from the scene tree, generating diverse instructions with LLM. We annotate over 2 million navigation instructions across 861 scenes and evaluate the data quality and navigation performance of trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。