提出首个针对机器人任务规划中模糊指代的基准,提升非专业用户使用体验。
REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
- 构建基于语用理论的模糊指代表达基准REI-Bench,系统评估指令模糊性影响。
- 发现模糊指代导致规划成功率下降最高达36.9%,多数失败因对象缺失。
- 提出面向任务的上下文认知方法,显著优于提示工程、思维链等现有策略。
机器人任务规划将人类指令分解为可执行的动作序列,使机器人完成复杂任务。尽管基于大语言模型(LLM)的规划器表现优异,但其假设人类指令清晰直接。现实中用户并非专家,其指令常含显著模糊性,尤其在老年人与儿童群体中更为普遍。语言学家指出,此类模糊性多源于指代表达(REs),其含义高度依赖对话上下文与环境。本文研究了指代表达模糊性对基于LLM的机器人任务规划的影响,并提出首个基于语用理论的机器人任务规划基准REI-Bench。实验发现,指代模糊性可使规划成功率下降高达36.9%,且多数失败源于规划器未能识别相关对象。为此,我们提出一种简单有效的解决方案:任务导向的上下文认知,生成更清晰的指令,性能超越提示感知、思维链及上下文学习等方法。该研究填补了对指令模糊性忽视的空白,推动真实场景任务规划发展,使机器人更易被非专业人士(如老人与儿童)使用。
原文摘要 · Abstract (English)
Robot task planning decomposes human instructions into executable action sequences that enable robots to complete a series of complex tasks. Although recent large language model (LLM)-based task planners achieve amazing performance, they assume that human instructions are clear and straightforward. However, real-world users are not experts, and their instructions to robots often contain significant vagueness. Linguists suggest that such vagueness frequently arises from referring expressions (REs), whose meanings depend heavily on dialogue context and environment. This vagueness is even more prevalent among the elderly and children, who are the groups that robots should serve more. This paper studies how such vagueness in REs within human instructions affects LLM-based robot task planning and how to overcome this issue. To this end, we propose the first robot task planning benchmark that systematically models vague REs grounded in pragmatic theory (REI-Bench), where we discover that the vagueness of REs can severely degrade robot planning performance, leading to success rate drops of up to 36.9%. We also observe that most failure cases stem from missing objects in planners. To mitigate the REs issue, we propose a simple yet effective approach: task-oriented context cognition, which generates clear instructions for robots, achieving state-of-the-art performance compared to aware prompts, chains of thought, and in-context learning. By tackling the overlooked issue of vagueness, this work contributes to the research community by advancing real-world task planning and making robots more accessible to non-expert users, e.g., the elderly and children.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。