arXiv:2601.02798cs.RO2026-01被引 1

用合成数据训练机器人,实现内镜导航不碰壁

Reinforcement Learning for Follow-the-Leader Robotic Endoscopic Navigation via Synthetic Data

  • 基于单目深度估计的强化学习框架,实现无接触导航
  • 深度精度提升39.2%,导航路径指数降低0.67
  • 适合医疗机器人、自主导航研究者参考

自主导航对医疗与工业内镜机器人至关重要,可实现狭窄管状环境的安全高效探索,避免与内壁接触是以往方法的长期挑战。本文提出一种基于柔性连续体结构的跟随领导者式内镜机器人,旨在最小化探头与肠道壁的接触,降低患者不适。为此,我们设计了一种基于单目深度估计的视觉强化学习框架。在NVIDIA Omniverse中构建了逼真的肠道仿真环境,用于训练与评估自主导航策略。同时,利用NVIDIA Replicator生成数千张合成腔内图像,微调Depth Anything模型,使单个单目摄像头即可实现肠道环境的密集三维感知。随后引入几何感知奖励与惩罚机制,实现精确管腔追踪。相较于原始Depth Anything模型,本方法深度精度δ₁提升39.2%;导航J-index相比第二佳方法降低0.67,验证了所提方法的鲁棒性与有效性。

原文摘要 · Abstract (English)

Autonomous navigation is crucial for both medical and industrial endoscopic robots, enabling safe and efficient exploration of narrow tubular environments without continuous human intervention, where avoiding contact with the inner walls has been a longstanding challenge for prior approaches. We present a follow-the-leader endoscopic robot based on a flexible continuum structure designed to minimize contact between the endoscope body and intestinal walls, thereby reducing patient discomfort. To achieve this objective, we propose a vision-based deep reinforcement learning framework guided by monocular depth estimation. A realistic intestinal simulation environment was constructed in \textit{NVIDIA Omniverse} to train and evaluate autonomous navigation strategies. Furthermore, thousands of synthetic intraluminal images were generated using NVIDIA Replicator to fine-tune the Depth Anything model, enabling dense three-dimensional perception of the intestinal environment with a single monocular camera. Subsequently, we introduce a geometry-aware reward and penalty mechanism to enable accurate lumen tracking. Compared with the original Depth Anything model, our method improves $δ_{1}$ depth accuracy by 39.2% and reduces the navigation J-index by 0.67 relative to the second-best method, demonstrating the robustness and effectiveness of the proposed approach.

内镜导航强化学习合成数据单目深度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。