arXiv:2506.04089cs.LGcs.AI2025-06ACL被引 10

构建厨房场景下可对比的模糊指令数据集,助力机器人理解用户意图。

AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment

  • 基于人类验证的文本数据集,包含1000组模糊与清晰指令对。
  • 涵盖三类模糊类型:人类偏好、常识知识、安全考量,共2000条任务。
  • 适合研究机器人对话理解、自然语言推理与任务规划的学者使用。

作为具身智能体的一部分,大语言模型(LLMs)通常根据用户提供的自然语言指令进行行为规划。然而,在真实环境中处理模糊指令仍是LLMs面临的挑战。尽管已有多种任务模糊性检测方法被提出,但由于测试数据集各异,缺乏统一基准,难以进行有效比较。为此,我们提出了AmbiK(Kitchen Environment中的模糊任务),一个完全以文本形式呈现的厨房场景下机器人指令模糊性数据集。AmbiK借助大语言模型辅助收集,并经人工验证,包含1000对模糊任务及其对应的无歧义版本,按模糊类型(人类偏好、常识知识、安全)分类,附带环境描述、澄清问题与答案、用户意图和任务计划,总计2000个任务。我们希望AmbiK能推动模糊性检测方法的统一评估。数据集已公开于https://github.com/cog-model/AmbiK-dataset。

原文摘要 · Abstract (English)

As a part of an embodied agent, Large Language Models (LLMs) are typically used for behavior planning given natural language instructions from the user. However, dealing with ambiguous instructions in real-world environments remains a challenge for LLMs. Various methods for task ambiguity detection have been proposed. However, it is difficult to compare them because they are tested on different datasets and there is no universal benchmark. For this reason, we propose AmbiK (Ambiguous Tasks in Kitchen Environment), the fully textual dataset of ambiguous instructions addressed to a robot in a kitchen environment. AmbiK was collected with the assistance of LLMs and is human-validated. It comprises 1000 pairs of ambiguous tasks and their unambiguous counterparts, categorized by ambiguity type (Human Preferences, Common Sense Knowledge, Safety), with environment descriptions, clarifying questions and answers, user intents, and task plans, for a total of 2000 tasks. We hope that AmbiK will enable researchers to perform a unified comparison of ambiguity detection methods. AmbiK is available at https://github.com/cog-model/AmbiK-dataset.

机器人自然语言模糊性数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。