arXiv:2501.04823cs.ROmath.OC2025-01被引 1

用少量人类反馈,让机器人自动识别潜在危险区域。

Learning Robot Safety from Sparse Human Feedback using Conformal Prediction

  • 基于最近邻分类与置信区间,从人类二值反馈中识别安全边界。
  • 在30次飞行测试中,显著提升模型预测控制器的安全性。
  • 适合需要高安全性的无人机、自动驾驶等场景使用。

确保机器人安全极具挑战:用户定义的约束可能遗漏边缘情况,即使训练数据安全,策略仍可能变得不安全,且安全标准具有主观性。为此,我们通过向人类展示策略轨迹并获取其对不安全行为的标记,利用统计方法中的置信预测(conformal prediction)识别出一个状态区域,该区域能以保证包含未来策略错误的指定比例(如95%)。该方法样本高效,基于最近邻分类,无需像传统置信预测那样预留数据。当机器人进入该可疑危险区时,系统可发出预警,模拟人类安全偏好并保证误报率。通过视频标注,系统可检测四轴飞行器视觉运动策略在穿越指定门框时的失败风险。我们提出一种通过避开可疑危险区来改进策略的方法,在6个导航任务中,30次四轴飞行实验验证了该方法对模型预测控制器安全性的提升。代码与视频已公开。

原文摘要 · Abstract (English)

Ensuring robot safety can be challenging; user-defined constraints can miss edge cases, policies can become unsafe even when trained from safe data, and safety can be subjective. Thus, we learn about robot safety by showing policy trajectories to a human who flags unsafe behavior. From this binary feedback, we use the statistical method of conformal prediction to identify a region of states, potentially in learned latent space, guaranteed to contain a user-specified fraction of future policy errors. Our method is sample-efficient, as it builds on nearest neighbor classification and avoids withholding data as is common with conformal prediction. By alerting if the robot reaches the suspected unsafe region, we obtain a warning system that mimics the human's safety preferences with guaranteed miss rate. From video labeling, our system can detect when a quadcopter visuomotor policy will fail to steer through a designated gate. We present an approach for policy improvement by avoiding the suspected unsafe region. With it we improve a model predictive controller's safety, as shown in experimental testing with 30 quadcopter flights across 6 navigation tasks. Code and videos are provided.

机器人安全置信预测强化学习无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。