让机器人在陌生环境中自动生成认知任务,真实评估其智能水平。
Automatic Cognitive Task Generation for In-Situ Evaluation of Embodied Agents
- 用图结构定义任务,分交互与演化两阶段动态生成新任务。
- 在10个未知场景中自动生成87,876个合理任务,覆盖日常认知能力。
- 揭示当前模型在3D交互和感知上严重不足,适合部署前评估使用。
随着通用智能体即将广泛部署于多样家庭环境,针对每个独特未见3D场景的评估成为关键前提。现有基准存在严重数据污染和场景特异性不足问题,难以评估代理在未知环境中的能力。为此,我们提出一种受人类认知启发的动态在位任务生成方法。通过结构化图表示定义任务,并构建两阶段交互-演化任务生成系统(TEA)。在交互阶段,代理主动与环境互动,形成任务执行与生成的闭环,实现持续生成;在演化阶段,任务图建模允许重组与复用已有任务,无需外部数据生成新任务。在10个未知场景的实验中,TEA在两轮内自动生成87,876个任务,经人工验证具备物理合理性并涵盖核心日常认知能力。在这些在位任务上对齐前沿模型与人类表现的基准测试显示,尽管模型在公开基准上表现优异,但在基础感知任务上表现意外不佳,严重缺乏3D交互意识,且推理对任务类型高度敏感。这些严峻发现凸显了在真实人机环境部署前进行在位评估的必要性。
原文摘要 · Abstract (English)
As general intelligent agents are poised for widespread deployment in diverse households, evaluation tailored to each unique unseen 3D environment has become a critical prerequisite. However, existing benchmarks suffer from severe data contamination and a lack of scene specificity, inadequate for assessing agent capabilities in unseen settings. To address this, we propose a dynamic in-situ task generation method for unseen environments inspired by human cognition. We define tasks through a structured graph representation and construct a two-stage interaction-evolution task generation system for embodied agents (TEA). In the interaction stage, the agent actively interacts with the environment, creating a loop between task execution and generation that allows for continuous task generation. In the evolution stage, task graph modeling allows us to recombine and reuse existing tasks to generate new ones without external data. Experiments across 10 unseen scenes demonstrate that TEA automatically generated 87,876 tasks in two cycles, which human verification confirmed to be physically reasonable and encompassing essential daily cognitive capabilities. Benchmarking SOTA models against humans on our in-situ tasks reveals that models, despite excelling on public benchmarks, perform surprisingly poorly on basic perception tasks, severely lack 3D interaction awareness and show high sensitivity to task types in reasoning. These sobering findings highlight the necessity of in-situ evaluation before deploying agents into real-world human environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。