无需末端传感器,通过混合强化学习实现手术机器人抓握的高精度力控。
Hybrid Offline-Online Reinforcement Learning for Sensorless, High-Precision Force Regulation in Surgical Robotic Grasping
- 构建物理一致的数字孪生模型,结合离线与在线强化学习。
- 仿真中力控误差小于1%,硬件实验平均误差低于4%。
- 适合追求无感测高精度力控的手术机器人研发人员。
在腱驱动手术器械中,精确抓握力调控受电机动力学、传动柔度、摩擦及远端机械特性间非线性耦合的根本限制。现有方法多依赖末端力传感或解析补偿,增加硬件复杂度或在动态运动下性能下降。本文提出一种无传感器控制框架,融合物理一致性建模与混合强化学习,在近端驱动的手术末端执行器上实现高精度远端力调控。我们建立达芬奇Xi抓取机构的第一性原理数字孪生模型,统一描述电、传动与钳口动力学。为安全学习该刚性且高度非线性的系统控制策略,提出三阶段流程:(i) 使用滚动时域CMA-ES生成动态可行的专家轨迹;(ii) 通过隐式Q学习进行完全离线策略学习,确保初始化稳定且无危险探索;(iii) 利用TD3进行在线优化,适应实际运行动态。所获策略直接将近端测量映射为电机电压,无需远端传感。仿真中,多谐波钳口运动下力控保持在设定值的1%以内;硬件实验显示多种轨迹下平均力误差低于4%,验证了从仿真到现实的迁移能力。学习策略约含71k参数,以kHz级速率执行,支持实时部署。结果表明,高保真建模与结构化离线-在线强化学习可实现无额外传感的精准远端力行为恢复,为手术机器人操作提供可扩展、机械兼容的解决方案。
原文摘要 · Abstract (English)
Precise grasp force regulation in tendon-driven surgical instruments is fundamentally limited by nonlinear coupling between motor dynamics, transmission compliance, friction, and distal mechanics. Existing solutions typically rely on distal force sensing or analytical compensation, increasing hardware complexity or degrading performance under dynamic motion. We present a sensorless control framework that combines physics-consistent modeling and hybrid reinforcement learning to achieve high-precision distal force regulation in a proximally actuated surgical end-effector. We develop a first-principles digital twin of the da Vinci Xi grasping mechanism that captures coupled electrical, transmission, and jaw dynamics within a unified differential-algebraic formulation. To safely learn control policies in this stiff and highly nonlinear system, we introduce a three-stage pipeline:(i)a receding-horizon CMA-ES oracle that generates dynamically feasible expert trajectories,(ii)fully offline policy learning via Implicit Q-Learning to ensure stable initialization without unsafe exploration, and (iii)online refinement using TD3 for adaptation to on-policy dynamics. The resulting policy directly maps proximal measurements to motor voltages and requires no distal sensing. In simulation, the controller maintains grasp force within 1% of the desired reference during multi-harmonic jaw motion. Hardware experiments demonstrate average force errors below 4% across diverse trajectories, validating sim-to-real transfer. The learned policy contains approximately 71k param and executes at kH rates, enabling real-time deployment. These results demonstrate that high-fidelity modeling combined with structured offline-online RL can recover precise distal force behavior without additional sensing, offering a scalable and mechanically compatible solution for surgical robotic manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。