让AI先猜人类意图,再执行指令,提升协作效率。
Infer Human's Intentions Before Following Natural Language Instructions
- 通过显式推理人类目标与意图,改进指令理解
- 在HandMeThat数据集上达当前最优性能
- 适合需要理解人类隐含意图的智能体场景
为了让人工智能代理在人类环境中协助完成日常合作任务,它们必须能理解自然语言指令。然而,真实人类指令天然具有模糊性,因为说话者假设对方已知晓其隐藏的目标与意图。标准的语言接地与规划方法无法解决此类模糊性,因其未将人类内在目标作为环境中的部分可观测因素建模。本文提出新的框架FISER(基于社会与具身推理的指令跟随),旨在提升协同具身任务中的自然语言指令遵循能力。该框架将对人类目标与意图的显式推断作为中间推理步骤。我们实现了一组基于Transformer的模型,并在挑战性基准HandMeThat上进行评估。实证表明,在制定行动方案前使用社会推理显式推断人类意图,优于纯端到端方法。与强基线(包括最大预训练语言模型的思维链提示)相比,FISER在所研究的具身社会推理任务中表现更优,在HandMeThat上达到当前最佳水平。
原文摘要 · Abstract (English)
For AI agents to be helpful to humans, they should be able to follow natural language instructions to complete everyday cooperative tasks in human environments. However, real human instructions inherently possess ambiguity, because the human speakers assume sufficient prior knowledge about their hidden goals and intentions. Standard language grounding and planning methods fail to address such ambiguities because they do not model human internal goals as additional partially observable factors in the environment. We propose a new framework, Follow Instructions with Social and Embodied Reasoning (FISER), aiming for better natural language instruction following in collaborative embodied tasks. Our framework makes explicit inferences about human goals and intentions as intermediate reasoning steps. We implement a set of Transformer-based models and evaluate them over a challenging benchmark, HandMeThat. We empirically demonstrate that using social reasoning to explicitly infer human intentions before making action plans surpasses purely end-to-end approaches. We also compare our implementation with strong baselines, including Chain of Thought prompting on the largest available pre-trained language models, and find that FISER provides better performance on the embodied social reasoning tasks under investigation, reaching the state-of-the-art on HandMeThat.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。