arXiv:2410.08852cs.ROcs.AI2024-10被引 12

让机器人在学动作时实时判断何时该问人,更准地应对专家行为变化。

Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback

  • 用间歇性反馈建模法,动态校准不确定度,保证预测区间覆盖率。
  • 在7自由度机械臂上测试,专家换策略时能及时发现高不确定性并主动求助。
  • 适合需要安全交互的机器人学习场景,尤其专家行为可能变化的部署环境。

在交互式模仿学习中,不确定性量化使机器人在部署时遭遇分布偏移时能主动向专家请求在线反馈。现有方法依赖集成分歧或蒙特卡洛丢弃,但在分布偏移下易产生过度自信估计。本文提出基于在线置信预测的间歇量化追踪(IQT)算法,利用间歇性人类标签的概率模型,在无分布假设下保持渐近覆盖性,并实现期望覆盖率。结合此方法,我们开发了ConformalDAgger,使机器人使用由IQT校准的预测区间作为部署时不确定性可靠指标,主动请求额外反馈。在模拟与真实硬件部署中,针对专家策略改变导致的分布偏移场景进行对比实验。结果表明,当专家行为发生变化时,ConformalDAgger能有效识别高不确定性,增加干预次数,显著加速机器人对新行为的学习。

原文摘要 · Abstract (English)

In interactive imitation learning (IL), uncertainty quantification offers a way for the learner (i.e. robot) to contend with distribution shifts encountered during deployment by actively seeking additional feedback from an expert (i.e. human) online. Prior works use mechanisms like ensemble disagreement or Monte Carlo dropout to quantify when black-box IL policies are uncertain; however, these approaches can lead to overconfident estimates when faced with deployment-time distribution shifts. Instead, we contend that we need uncertainty quantification algorithms that can leverage the expert human feedback received during deployment time to adapt the robot's uncertainty online. To tackle this, we draw upon online conformal prediction, a distribution-free method for constructing prediction intervals online given a stream of ground-truth labels. Human labels, however, are intermittent in the interactive IL setting. Thus, from the conformal prediction side, we introduce a novel uncertainty quantification algorithm called intermittent quantile tracking (IQT) that leverages a probabilistic model of intermittent labels, maintains asymptotic coverage guarantees, and empirically achieves desired coverage levels. From the interactive IL side, we develop ConformalDAgger, a new approach wherein the robot uses prediction intervals calibrated by IQT as a reliable measure of deployment-time uncertainty to actively query for more expert feedback. We compare ConformalDAgger to prior uncertainty-aware DAgger methods in scenarios where the distribution shift is (and isn't) present because of changes in the expert's policy. We find that in simulated and hardware deployments on a 7DOF robotic manipulator, ConformalDAgger detects high uncertainty when the expert shifts and increases the number of interventions compared to baselines, allowing the robot to more quickly learn the new behavior.

模仿学习不确定性量化机器人交互在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。