arXiv:2409.12153cs.RO2024-09ICRA被引 9

让机器人安全地利用影响力,提升协作效率。

Robots that Learn to Safely Influence via Prediction-Informed Reach-Avoid Dynamic Games

  • 构建预测驱动的动态博弈模型,融合人机交互意图与不确定性。
  • 在39维仿真任务中实现高安全率下机器人主动影响人类行为。
  • 适合需要高效人机协作且重视安全的机器人系统设计者。

机器人可通过影响人类行为更高效完成任务:自动驾驶汽车可在交叉口轻微前移以通过,桌面机械臂可优先取物。然而,若盲目执行,这种影响可能危及附近人员安全。本文提出并求解一种新型鲁棒性‘到达-回避’动态博弈,使机器人仅在存在安全备用控制时才最大化其影响力。人类行为建模为以目标为导向但受机器人计划影响,从而捕捉影响力。机器人侧在物理与信念联合空间求解动态博弈,可推理其对人类行为的不确定性如何随时间演变。我们实例化该方法(称为SLIDE,Safely Leveraging Influence in Dynamic Environments),在39维模拟人机协作操作任务中通过离线博弈强化学习求解。与将人类视为最坏情况对手的鲁棒基线、未显式建模影响力的安全部控制器以及基于能量函数的安全防护器相比,SLIDE始终能在安全前提下利用机器人对人类的影响,使机器人表现更积极,同时保障高安全率的任务执行。

原文摘要 · Abstract (English)

Robots can influence people to accomplish their tasks more efficiently: autonomous cars can inch forward at an intersection to pass through, and tabletop manipulators can go for an object on the table first. However, a robot's ability to influence can also compromise the safety of nearby people if naively executed. In this work, we pose and solve a novel robust reach-avoid dynamic game which enables robots to be maximally influential, but only when a safety backup control exists. On the human side, we model the human's behavior as goal-driven but conditioned on the robot's plan, enabling us to capture influence. On the robot side, we solve the dynamic game in the joint physical and belief space, enabling the robot to reason about how its uncertainty in human behavior will evolve over time. We instantiate our method, called SLIDE (Safely Leveraging Influence in Dynamic Environments), in a high-dimensional (39-D) simulated human-robot collaborative manipulation task solved via offline game-theoretic reinforcement learning. We compare our approach to a robust baseline that treats the human as a worst-case adversary, a safety controller that does not explicitly reason about influence, and an energy-function-based safety shield. We find that SLIDE consistently enables the robot to leverage the influence it has on the human when it is safe to do so, ultimately allowing the robot to be less conservative while still ensuring a high safety rate during task execution.

人机协作动态博弈安全控制机器人决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。