arXiv:2608.29023cs.RO2026-08

用错误示范教人理解机器人行为,提升长期记忆效果

Teaching Robot Policies to Humans Using Erroneous Examples

论文配图:Teaching Robot Policies to Humans Using Erroneous Examples
图 1 · 摘自论文原文
  • 通过展示错误的机器人行为并让用户纠正,模拟教学中的错例学习
  • 用户在修正错误时口头解释推理过程,显著提升政策记忆保持率
  • 发现逆强化学习式思考者在预测任务中表现最佳,适合个性化教学

人机协作指人类与自主代理共同完成目标的过程。当机器人策略(即不同情境下的行为)对人类透明时,协作效果最佳。现有研究多采用演示式解释,并借鉴教育学成果改进人类对机器人策略的学习方式。然而,尚无单一方法在不同领域、难度、学习者间均有效,如何最有效地教授机器人策略仍待解决。传统课堂中,学生通过反思和纠正错误答案来掌握知识盲点。本文提出将错误示例用于教授机器人策略,扩展现有策略教学框架。我们开展用户研究,参与者观看机器人行为的错误示范,并纠正动作以匹配真实策略。结果表明,观看错误示范并口头阐述预测逻辑能有效提升策略长期记忆,与课堂错例学习效应一致。我们还根据学习风格分类,发现采用逆强化学习式推理的参与者在策略预测任务中表现最优。本工作旨在推进机器人向人类传授其策略的方法。

原文摘要 · Abstract (English)

Human-robot collaboration describes the process of humans and autonomous agents working together to accomplish common goals. This process is facilitated best when robot policies, or behaviors in different situations, are made transparent to humans. Demonstration-based explanations have been a focus of human-robot collaboration research, and the field has frequently drawn upon literature from education to improve how humans are taught robot policies. However, no single teaching method has been proven effective across domains, difficulties, learners, and other variables; the question of how humans can most effectively be taught robot policies remains open. In traditional classrooms, learners are shown erroneous examples, in which they reflect on and correct incorrect responses to understand common pitfalls when learning a concept. We propose using erroneous examples to teach robot policies, extending an existing policy teaching framework. We conduct a user study in which participants view incorrect demonstrations of robot behavior and correct the actions to align with the actual policy. Our findings suggest that viewing these incorrect demonstrations and verbalizing one's reasoning in predicting a robot's actions improves retention of the policy over time, in agreement with the effect of erroneous examples in classrooms. We also categorize participants into distinct learning styles and establish that participants using inverse reinforcement learning-like reasoning perform best on policy prediction tasks. With this work, we aim to advance the methods by which robots educate humans on their policies.

人机协作错误示例策略教学学习风格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。