让手术机器人像人一样持续学习新技能,边学边用。
SurgIRL: Towards Life-Long Learning for Surgical Automation by Incremental Reinforcement Learning
- 用增量强化学习整合已有手术策略,逐步积累能力。
- 在仿真中高效完成10项任务的独立或连续学习。
- 支持从仿真到真实机器人(dVRK)的可靠迁移。
手术自动化有望显著提升手术效果和可及性。近期研究采用强化学习训练独立策略以自动化各类手术任务,但这些策略互不共享,任务变化时需重新学习,效率低下。受人类外科医生积累经验的启发,我们提出外科增量强化学习(SurgIRL),旨在(1)通过参考外部策略(知识)获取新技能,(2)累积并复用这些技能,实现多个未见任务的渐进式学习。SurgIRL框架包含三大组件:首先构建一个可扩展的知识集,包含异构且对任务有帮助的策略;其次提出知识包容注意力网络与最大覆盖探索(KIAN-ACE),通过最大化知识集覆盖提升探索效率;最后基于KIAN-ACE设计增量学习流水线,实现知识的积累与复用。仿真实验表明,KIAN-ACE能高效完成10项手术任务的单独或增量学习。我们在da Vinci Research Kit(dVRK)上评估了所学策略,成功实现从仿真到真实的迁移。
原文摘要 · Abstract (English)
Surgical automation holds immense potential to improve the outcome and accessibility of surgery. Recent studies use reinforcement learning to learn policies that automate different surgical tasks. However, these policies are developed independently and are limited in their reusability when the task changes, making it more time-consuming when robots learn to solve multiple tasks. Inspired by how human surgeons build their expertise, we train surgical automation policies through Surgical Incremental Reinforcement Learning (SurgIRL). SurgIRL aims to (1) acquire new skills by referring to external policies (knowledge) and (2) accumulate and reuse these skills to solve multiple unseen tasks incrementally (incremental learning). Our SurgIRL framework includes three major components. We first define an expandable knowledge set containing heterogeneous policies that can be helpful for surgical tasks. Then, we propose Knowledge Inclusive Attention Network with mAximum Coverage Exploration (KIAN-ACE), which improves learning efficiency by maximizing the coverage of the knowledge set during the exploration process. Finally, we develop incremental learning pipelines based on KIAN-ACE to accumulate and reuse learned knowledge and solve multiple surgical tasks sequentially. Our simulation experiments show that KIAN-ACE efficiently learns to automate ten surgical tasks separately or incrementally. We also evaluate our learned policies on the da Vinci Research Kit (dVRK) and demonstrate successful sim-to-real transfers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。