arXiv:2410.18825cs.RO2024-10被引 3

提出一套通用框架,让机器人系统在K8s上自动应对故障并保持状态。

A generic approach for reactive stateful mitigation of application failures in distributed robotics systems deployed with Kubernetes

  • 用行为树实现可复用的故障监控与响应策略。
  • 在模拟环境中验证了自主导航和机械臂操作的故障恢复能力。
  • 适合部署在K8s上的复杂机器人应用,尤其关注状态一致性。

将计算密集型算法卸载到边缘或云端,为解决机器人系统车载算力与能耗限制提供了有效途径。在基于容器管理平台 Kubernetes(K8s)部署的云原生应用中,确保对各类故障的弹性至关重要。然而,与物理世界交互的复杂机器人系统带来了云原生领域尚未覆盖的特定挑战与需求。本文提出一种新型方法,用于在 Kubernetes 上部署的分布式机器人系统中实现机器人系统监控及有状态、反应式的故障缓解,结合机器人操作系统 ROS2 与行为树(Behaviour Trees)的通用架构,支持任意复杂的监控与容错策略。该方法在两个示例应用中得到验证:自主移动机器人(AMR)导航与模拟环境中的机器人抓取操作,证明其有效性与应用无关性。

原文摘要 · Abstract (English)

Offloading computationally expensive algorithms to the edge or even cloud offers an attractive option to tackle limitations regarding on-board computational and energy resources of robotic systems. In cloud-native applications deployed with the container management system Kubernetes (K8s), one key problem is ensuring resilience against various types of failures. However, complex robotic systems interacting with the physical world pose a very specific set of challenges and requirements that are not yet covered by failure mitigation approaches from the cloud-native domain. In this paper, we therefore propose a novel approach for robotic system monitoring and stateful, reactive failure mitigation for distributed robotic systems deployed using Kubernetes (K8s) and the Robot Operating System (ROS2). By employing the generic substrate of Behaviour Trees, our approach can be applied to any robotic workload and supports arbitrarily complex monitoring and failure mitigation strategies. We demonstrate the effectiveness and application-agnosticism of our approach on two example applications, namely Autonomous Mobile Robot (AMR) navigation and robotic manipulation in a simulated environment.

机器人系统Kubernetes容错机制ROS2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。