用扩散模型训练手术机器人,能从失败数据中学习并保持稳定
Diffusion Stabilizer Policy for Automated Surgical Robot Manipulations
- 基于扩散模型构建双阶段策略,先用干净数据训练,再融合扰动数据持续优化
- 在无扰动场景下性能超越基准,在含扰动演示下仍保持高鲁棒性
- 适合缺乏高质量示范数据的手术机器人自动化研究者
智能手术机器人有望通过更精准、自动化的手术流程彻底改变临床实践。然而,与家庭操作任务的进展相比,手术机器人的自动化仍处于探索阶段。近期的成功主要依赖于(1)先进的模型(如Transformer和扩散模型)和(2)大规模数据利用。为将这些成果拓展至手术机器人领域,我们提出一种基于扩散模型的策略学习框架——扩散稳定器策略(DSP),支持使用不完美甚至失败的轨迹进行训练。该方法包含两个阶段:首先仅使用干净数据训练扩散稳定器策略;随后,通过混合干净与扰动数据,并依据动作预测误差进行筛选,持续更新策略。在多种手术环境中的全面实验表明,该方法在无扰动条件下表现优异,且在处理扰动示范时展现出强鲁棒性。
原文摘要 · Abstract (English)
Intelligent surgical robots have the potential to revolutionize clinical practice by enabling more precise and automated surgical procedures. However, the automation of such robot for surgical tasks remains under-explored compared to recent advancements in solving household manipulation tasks. These successes have been largely driven by (1) advanced models, such as transformers and diffusion models, and (2) large-scale data utilization. Aiming to extend these successes to the domain of surgical robotics, we propose a diffusion-based policy learning framework, called Diffusion Stabilizer Policy (DSP), which enables training with imperfect or even failed trajectories. Our approach consists of two stages: first, we train the diffusion stabilizer policy using only clean data. Then, the policy is continuously updated using a mixture of clean and perturbed data, with filtering based on the prediction error on actions. Comprehensive experiments conducted in various surgical environments demonstrate the superior performance of our method in perturbation-free settings and its robustness when handling perturbed demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。