用自然语言定义机器人任务,让普通人也能参与评估。
RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains
- 用户用自然语言编写可执行的机器人任务,系统自动生成规范指令。
- 语言定义的任务家族暴露了传统基准下无法发现的泛化失败。
- 多人共创任务能持续扩展评估空间,提升多样性与可比性。
机器人操作系统的评估长期依赖少数专家制定的固定基准,任务实例、约束和成功标准预先设定且难以扩展。这一模式限制了评估设计的参与者,并模糊了策略对用户自定义任务意图、约束和成功定义的变化响应。我们提出将评估重构为基于语言的结构化物理域上的过程。本文介绍RoboPlayground框架,允许用户在结构化物理环境中使用自然语言编写可执行的操作任务。自然语言指令被编译为包含显式资产定义、初始化分布和成功谓词的可复现任务规范。每条指令定义一组相关任务,支持受控的语义与行为变化,同时保持可执行性和可比性。我们在结构化积木操作领域实现了该框架,并从三个维度进行评估:用户研究显示,语言界面比编程基线更易用、认知负荷更低;在语言定义的任务族上评估学习策略,揭示了固定基准下未察觉的泛化失败;最后,我们证明任务多样性随贡献者多样性增长,而非仅依赖任务数量,从而实现评估空间的持续扩展。项目主页:https://roboplayground.github.io
原文摘要 · Abstract (English)
Evaluation of robotic manipulation systems has largely relied on fixed benchmarks authored by a small number of experts, where task instances, constraints, and success criteria are predefined and difficult to extend. This paradigm limits who can shape evaluation and obscures how policies respond to user-authored variations in task intent, constraints, and notions of success. We argue that evaluating modern manipulation policies requires reframing evaluation as a language-driven process over structured physical domains. We present RoboPlayground, a framework that enables users to author executable manipulation tasks using natural language within a structured physical domain. Natural language instructions are compiled into reproducible task specifications with explicit asset definitions, initialization distributions, and success predicates. Each instruction defines a structured family of related tasks, enabling controlled semantic and behavioral variation while preserving executability and comparability. We instantiate RoboPlayground in a structured block manipulation domain and evaluate it along three axes. A user study shows that the language-driven interface is easier to use and imposes lower cognitive workload than programming-based and code-assist baselines. Evaluating learned policies on language-defined task families reveals generalization failures that are not apparent under fixed benchmark evaluations. Finally, we show that task diversity scales with contributor diversity rather than task count alone, enabling evaluation spaces to grow continuously through crowd-authored contributions. Project Page: https://roboplayground.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。