让机器人导航指令更丰富,基于环境结构与语义生成。
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
- 结合环境结构与语义知识生成导航指令
- 用对抗式奖励学习提升指令多样性与质量
- 适合研究具身智能与自然语言生成的学者
具身人工智能致力于开发能理解并执行人类语言指令,同时以自然语言进行交流的机器人。本文研究为具身机器人生成高度详细的导航指令任务。尽管近期研究在从图像序列生成分步指令方面取得显著进展,但生成的指令在物体和地标指代上缺乏多样性。现有说话者模型会学习规避评估指标,获得高分却生成低质量句子。为此,本文提出SAS(空间感知说话者)模型,利用环境的结构与语义知识生成更丰富的指令。训练中采用对抗式奖励学习方法,避免语言评估指标带来的系统性偏差。实验证明,该方法在标准指标下优于现有指令生成模型。代码已开源。
原文摘要 · Abstract (English)
Embodied AI aims to develop robots that can \textit{understand} and execute human language instructions, as well as communicate in natural languages. On this front, we study the task of generating highly detailed navigational instructions for the embodied robots to follow. Although recent studies have demonstrated significant leaps in the generation of step-by-step instructions from sequences of images, the generated instructions lack variety in terms of their referral to objects and landmarks. Existing speaker models learn strategies to evade the evaluation metrics and obtain higher scores even for low-quality sentences. In this work, we propose SAS (Spatially-Aware Speaker), an instruction generator or \textit{Speaker} model that utilises both structural and semantic knowledge of the environment to produce richer instructions. For training, we employ a reward learning method in an adversarial setting to avoid systematic bias introduced by language evaluation metrics. Empirically, our method outperforms existing instruction generation models, evaluated using standard metrics. Our code is available at \url{https://github.com/gmuraleekrishna/SAS}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。