arXiv:2409.02724cs.RO2024-09被引 1

仅用状态数据实现手术机器人自动化,突破动作标签难获取的瓶颈。

Surgical Task Automation Using Actor-Critic Frameworks and Self-Supervised Imitation Learning

  • 通过自监督模仿学习,从纯状态演示中提取动作信息
  • 在仿真平台上性能接近依赖动作标签的方法
  • 适合缺乏动作标注但有专家操作视频的医疗AI场景

手术机器人任务自动化近年来备受关注,有望同时提升外科医生与患者体验。基于强化学习(RL)的方法在多种手术任务中展现出自动化操控的潜力。为应对探索难题,可利用专家示范提升学习效率,但现有模仿学习方法通常需同时具备状态和动作标签。然而动作标签难以捕捉,且人工标注成本高昂。因此,如何利用仅含状态的专家示范进行强化学习仍是开放性挑战。本文提出一种名为AC-SSIL的演员-评论家框架,通过自监督模仿学习(SSIL)方法,从查询状态的最近邻中检索动作信息,并结合演员网络的自举机制,将纯状态示范融入强化学习范式。在开源手术仿真平台上的实验表明,该方法显著优于基础强化学习模型,且性能与依赖动作标签的模仿学习方法相当,验证了其在专家示范引导学习场景中的有效性和潜力。

原文摘要 · Abstract (English)

Surgical robot task automation has recently attracted great attention due to its potential to benefit both surgeons and patients. Reinforcement learning (RL) based approaches have demonstrated promising ability to provide solutions to automated surgical manipulations on various tasks. To address the exploration challenge, expert demonstrations can be utilized to enhance the learning efficiency via imitation learning (IL) approaches. However, the successes of such methods normally rely on both states and action labels. Unfortunately action labels can be hard to capture or their manual annotation is prohibitively expensive owing to the requirement for expert knowledge. It therefore remains an appealing and open problem to leverage expert demonstrations composed of pure states in RL. In this work, we present an actor-critic RL framework, termed AC-SSIL, to overcome this challenge of learning with state-only demonstrations collected by following an unknown expert policy. It adopts a self-supervised IL method, dubbed SSIL, to effectively incorporate demonstrated states into RL paradigms by retrieving from demonstrates the nearest neighbours of the query state and utilizing the bootstrapping of actor networks. We showcase through experiments on an open-source surgical simulation platform that our method delivers remarkable improvements over the RL baseline and exhibits comparable performance against action based IL methods, which implies the efficacy and potential of our method for expert demonstration-guided learning scenarios.

手术自动化强化学习模仿学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。