arXiv:2509.00329cs.ROcs.AI2025-09被引 1

针对可变形机器人的动态导航,提出分阶段的强化学习方法,提升规划效率与泛化能力。

Jacobian Exploratory Dual-Phase Reinforcement Learning for Dynamic Endoluminal Navigation of Deformable Continuum Robots

  • 分两阶段:先小范围探索估计变形雅可比矩阵,再用其增强状态表示
  • 相比PPO基准,收敛快3.2倍,达目标少25%步数,未知环境下成功率高33%
  • 适合需要高适应性与实时性的内镜手术机器人控制场景

可变形连续体机器人(DCRs)因非线性形变力学和部分可观测性,违背传统强化学习的马尔可夫假设,带来独特规划挑战。尽管基于雅可比的方法在刚性机械臂上有理论基础,但其在时间变化的运动学与欠驱动形变动力学下的应用受限。本文提出雅可比探索式双阶段强化学习(JEDP-RL),将规划分解为雅可比估计与策略执行两个阶段。每个训练步骤中,先通过小规模局部探索动作估计形变雅可比矩阵,再将其作为特征融入状态表示,以恢复近似马尔可夫性。基于SOFA手术动态仿真的大量实验表明,相较于近端策略优化(PPO)基线,JEDP-RL具有三大优势:1)收敛速度提升3.2倍;2)导航效率提高,到达目标所需步数减少25%;3)泛化能力强,在材料属性变化下取得92%成功率,在未见过的组织环境中达到83%成功率(比PPO高33个百分点)。

原文摘要 · Abstract (English)

Deformable continuum robots (DCRs) present unique planning challenges due to nonlinear deformation mechanics and partial state observability, violating the Markov assumptions of conventional reinforcement learning (RL) methods. While Jacobian-based approaches offer theoretical foundations for rigid manipulators, their direct application to DCRs remains limited by time-varying kinematics and underactuated deformation dynamics. This paper proposes Jacobian Exploratory Dual-Phase RL (JEDP-RL), a framework that decomposes planning into phased Jacobian estimation and policy execution. During each training step, we first perform small-scale local exploratory actions to estimate the deformation Jacobian matrix, then augment the state representation with Jacobian features to restore approximate Markovianity. Extensive SOFA surgical dynamic simulations demonstrate JEDP-RL's three key advantages over proximal policy optimization (PPO) baselines: 1) Convergence speed: 3.2x faster policy convergence, 2) Navigation efficiency: requires 25% fewer steps to reach the target, and 3) Generalization ability: achieve 92% success rate under material property variations and achieve 83% (33% higher than PPO) success rate in the unseen tissue environment.

机器人控制强化学习可变形机器人内镜导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。