arXiv:2505.19574cs.ROcs.AI2025-05被引 1

让机器人实时识别环境变化,自适应调整行为,提升复杂地形下的安全性和效率。

Situationally-Aware Dynamics Learning

  • 通过联合建模状态转移,在线学习隐藏状态分布以识别运行情境。
  • 在仿真与真实场景中实现数据效率提升,策略性能显著改善。
  • 适合需在不确定环境中自主决策的机器人系统,如野外巡检、救援导航。

在复杂非结构化环境中,自主机器人因隐含的未观测因素而难以理解自身状态与外部世界。为此,我们提出一种新型在线学习框架,用于隐状态表示的学习,使机器人能够实时适应不确定且动态变化的条件,避免模糊判断和错误行为。该方法形式化为广义隐藏参数马尔可夫决策过程,显式建模未观测参数对状态转移和奖励结构的影响。核心创新在于在线学习状态转移的联合分布,作为潜在自我与环境因素的表达表征。该概率方法支持对不同运行情境的识别与适应,增强鲁棒性与安全性。通过多变量贝叶斯在线变点检测的扩展,本方法可分割驱动机器人动力学的底层数据生成过程的变化。基于最新状态转移的联合分布,生成当前情境的符号化表示,从而指导自适应、情境感知的决策。在非结构化地形导航任务中验证了其有效性,未建模的地形特性会显著影响运动表现。大量仿真与实测实验表明,该方法在数据效率、策略性能上均有显著提升,并涌现出更安全、适应性强的导航策略。

原文摘要 · Abstract (English)

Autonomous robots operating in complex, unstructured environments face significant challenges due to latent, unobserved factors that obscure their understanding of both their internal state and the external world. Addressing this challenge would enable robots to develop a more profound grasp of their operational context. To tackle this, we propose a novel framework for online learning of hidden state representations, with which the robots can adapt in real-time to uncertain and dynamic conditions that would otherwise be ambiguous and result in suboptimal or erroneous behaviors. Our approach is formalized as a Generalized Hidden Parameter Markov Decision Process, which explicitly models the influence of unobserved parameters on both transition dynamics and reward structures. Our core innovation lies in learning online the joint distribution of state transitions, which serves as an expressive representation of latent ego- and environmental-factors. This probabilistic approach supports the identification and adaptation to different operational situations, improving robustness and safety. Through a multivariate extension of Bayesian Online Changepoint Detection, our method segments changes in the underlying data generating process governing the robot's dynamics. The robot's transition model is then informed with a symbolic representation of the current situation derived from the joint distribution of latest state transitions, enabling adaptive and context-aware decision-making. To showcase the real-world effectiveness, we validate our approach in the challenging task of unstructured terrain navigation, where unmodeled and unmeasured terrain characteristics can significantly impact the robot's motion. Extensive experiments in both simulation and real world reveal significant improvements in data efficiency, policy performance, and the emergence of safer, adaptive navigation strategies.

机器人学习自适应控制状态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。