用人类干预提升机器人胃镜导航安全,避免碰撞。
Safe Navigation for Robotic Digestive Endoscopy via Human Intervention-based Reinforcement Learning
- 引入人类干预机制优化强化学习策略,提升探索效率。
- 在模拟中实现8.02mm平均轨迹误差和0.862安全得分。
- 适合医疗机器人、智能内窥镜系统研究者参考。
随着自动化机器人消化内窥镜(RDE)应用日益广泛,如何在结构复杂且狭窄的消化道中实现安全高效的导航成为关键挑战。现有自动化强化学习导航算法因缺乏必要的人类干预,常导致潜在碰撞,严重限制了其在临床实践中的安全性与有效性。为此,本文提出一种基于人类干预的近端策略优化框架(HI-PPO),融合增强探索机制(EEM)、奖励-惩罚调节(RPA)和行为克隆相似性(BCS),以解决PPO在复杂胃肠道环境中探索效率低的问题。在仿真平台上进行对比实验,结果表明,HI-PPO实现了8.02 mm的平均轨迹误差(ATE)和0.862的安全得分,性能接近人类专家水平。代码将在论文发表后公开。
原文摘要 · Abstract (English)
With the increasing application of automated robotic digestive endoscopy (RDE), ensuring safe and efficient navigation in the unstructured and narrow digestive tract has become a critical challenge. Existing automated reinforcement learning navigation algorithms often result in potentially risky collisions due to the absence of essential human intervention, which significantly limits the safety and effectiveness of RDE in actual clinical practice. To address this limitation, we proposed a Human Intervention (HI)-based Proximal Policy Optimization (PPO) framework, dubbed HI-PPO, which incorporates expert knowledge to enhance RDE's safety. Specifically, HI-PPO combines Enhanced Exploration Mechanism (EEM), Reward-Penalty Adjustment (RPA), and Behavior Cloning Similarity (BCS) to address PPO's exploration inefficiencies for safe navigation in complex gastrointestinal environments. Comparative experiments were conducted on a simulation platform, and the results showed that HI-PPO achieved a mean ATE (Average Trajectory Error) of \(8.02\ \text{mm}\) and a Security Score of \(0.862\), demonstrating performance comparable to human experts. The code will be publicly available once this paper is published.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。