arXiv:2410.00774cs.ROcs.AI2024-10

通过预测未来不确定性,让机器人在未知环境中自适应调整动作。

Adaptive Motion Generation Using Uncertainty-Driven Foresight Prediction

  • 用动态内仿真预测多种未来,选不确定性最低的路径优化控制
  • 在门开任务中成功区分推、拉、滑三种方式,传统方法失败
  • 首次将不确定性嵌入策略而非观测,适合复杂交互场景

环境不确定性是实世界机器人任务中的长期难题,因意外观测无法通过人工脚本覆盖。基于学习的控制方法虽具灵活性,但受限于确定性本质,在不确定性下仍易失效。为实现自适应执行目标任务,控制模型需准确理解潜在不确定性,并探索最优动作以最小化不确定性。本文扩展了一种基于预测学习的机器人控制方法,采用动态内仿真进行前视预测:通过采样多个可能未来,替换隐藏状态为导致未来不确定性更低的版本。在门开任务中评估模型表现,该门可通过推、拉或滑三种方式开启,机器人无法视觉区分,需实时自适应。结果表明,所提模型能通过与门的交互自适应分化动作,而传统方法无法稳定分化。通过对RNN隐藏状态的李雅普诺夫指数分析发现,前视模块使模型关注未来后果,将不确定性嵌入策略而非结果观测,有利于生成探索性多样化动作。

原文摘要 · Abstract (English)

Uncertainty of environments has long been a difficult characteristic to handle, when performing real-world robot tasks. This is because the uncertainty produces unexpected observations that cannot be covered by manual scripting. Learning based robot controlling methods are a promising approach for generating flexible motions against unknown situations, but still tend to suffer under uncertainty due to its deterministic nature. In order to adaptively perform the target task under such conditions, the robot control model must be able to accurately understand the possible uncertainty, and to exploratively derive the optimal action that minimizes such uncertainty. This paper extended an existing predictive learning based robot control method, which employ foresight prediction using dynamic internal simulation. The foresight module refines the model's hidden states by sampling multiple possible futures and replace with the one that led to the lower future uncertainty. The adaptiveness of the model was evaluated on a door opening task. The door can be opened either by pushing, pulling, or sliding, but robot cannot visually distinguish which way, and is required to adapt on the fly. The results showed that the proposed model adaptively diverged its motion through interaction with the door, whereas conventional methods failed to stably diverge. The models were analyzed on Lyapunov exponents of RNN hidden states which reflect the possible divergence at each time step during task execution. The result indicated that the foresight module biased the model to consider future consequences, which lead to embedding uncertainties at the policy of the robot controller, rather than the resultant observation. This is beneficial for implementing adaptive behaviors, which indices derivation of diverse motion during exploration.

机器人控制不确定性建模前视预测自适应运动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。