arXiv:2605.08269cs.ROcs.SY2026-05

用解剖标志点引导强化学习,实现胃内自主导航。

Anatomical Landmark-Guided Deep Reinforcement Learning for Autonomous Gastric Navigation

论文配图:Anatomical Landmark-Guided Deep Reinforcement Learning for Autonomous Gastric Navigation
图 1 · 摘自论文原文
  • 基于解剖标志点坐标而非视频流,降低控制复杂度
  • 仿真中覆盖率达97%以上,50秒内完成导航
  • 适合胃镜胶囊机器人自主导航研究者参考

无线胶囊内窥镜(WCE)可无痛观察消化道,但其诊断效果受限于黏膜覆盖不全以及现有导航方法在不同患者间迁移性差。本文提出一种可迁移的解剖标志点引导深度强化学习(AL-DRL)框架,用于自主胃内导航。通过轻量级边缘-轮廓-深度融合模块,策略基于稳定、低维的标志点坐标运行,而非高维视频流,有效缩小仿真到现实的差距。在8个基于患者模型的仿真中,该方法在50秒内实现超过97%的覆盖率,显著优于原始PPO、SAC和DQN智能体。采用两阶段仿真到现实的迁移流程,结合自适应动态规划控制器,主动缓解物理扰动。离体实验显示,平均覆盖率可达87%,较专家手动控制减少53%的操作时间。

原文摘要 · Abstract (English)

Wireless capsule endoscopy (WCE) enables painless visualization of the gastrointestinal tract, but its diagnostic potential is limited by incomplete mucosal coverage and poor transferability of existing navigation methods across patient anatomies. We propose a transferable, anatomical landmarkguided deep reinforcement learning (AL-DRL) framework for autonomous gastric navigation. Leveraging a lightweight edgecontour-depth fusion module, our policy operates on stable, lowdimensional landmark coordinates rather than high-dimensional video streams, effectively bridging the sim-to-real gap. In simulations across eight patient-derived models, the method achieves over 97% coverage within 50 seconds, significantly outperforming vanilla PPO, SAC, and DQN agents. A two-stage sim-to-real pipeline with an adaptive dynamic programming controller actively mitigates physical disturbances. Ex-vivo experiments demonstrate a mean coverage of 87% and a 53% reduction in procedure time compared with expert manual control.

自主导航强化学习胶囊内镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。