人类示范中的非最优行为具有系统性,可分类为四类路径偏移。
Demonstration Sidetracks: Categorizing Systematic Non-Optimality in Human Demonstrations
- 识别出四种系统性示范偏差:探索、错误、对齐、暂停
- 40人实验中偏差普遍存在,分布与任务上下文相关
- 不同控制界面影响用户行为模式,适用于机器人学习研究
从示范学习(LfD)是机器人习得新技能的常用方法,但多数方法将人类示范的不完美视为随机噪声。本文研究非专家示范中的非最优行为,发现其具有系统性,称之为示范路径偏移(sidetracks)。通过40名参与者在真实空间中完成长时程机器人任务的实验,重建场景至仿真环境并标注所有示范。识别出四类偏移类型(探索、错误、对齐、暂停)及一种控制基准模式(一维控制)。偏移在各参与者中频繁出现,其时空分布与任务上下文密切相关。此外,用户控制模式受控制界面影响。这些发现表明需改进对非最优示范的建模,以提升LfD算法性能,弥合实验室训练与现实部署的差距。所有示范数据、基础设施与标注已开源于https://github.com/AABL-Lab/Human-Demonstration-Sidetracks。
原文摘要 · Abstract (English)
Learning from Demonstration (LfD) is a popular approach for robots to acquire new skills, but most LfD methods suffer from imperfections in human demonstrations. Prior work typically treats these suboptimalities as random noise. In this paper we study non-optimal behaviors in non-expert demonstrations and show that they are systematic, forming what we call demonstration sidetracks. Using a public space study with 40 participants performing a long-horizon robot task, we recreated the setup in simulation and annotated all demonstrations. We identify four types of sidetracks (Exploration, Mistake, Alignment, Pause) and one control pattern (one-dimension control). Sidetracks appear frequently across participants, and their temporal and spatial distribution is tied to task context. We also find that users' control patterns depend on the control interface. These insights point to the need for better models of suboptimal demonstrations to improve LfD algorithms and bridge the gap between lab training and real-world deployment. All demonstrations, infrastructure, and annotations are available at https://github.com/AABL-Lab/Human-Demonstration-Sidetracks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。