用视觉语言动作模型实现胆道镜自主导航,提升操作精度与安全性。
BiliVLA: Scene-Aware Vision-Language-Action Model with Reinforcement Learning for Autonomous Biliary Endoscopic Navigation

- 将导航任务建模为指令驱动的视觉运动学习,融合语义定位与动作预测。
- 在实体模型上实现91.96%动作准确率、84.85%成功率和0.9625平均交并比。
- 通过场景感知监督和安全恢复机制,增强对复杂解剖结构的适应能力。
内镜逆行胰胆管造影(ERCP)要求在狭窄单目视野中实现精准内镜导航和稳定的胆管插管,该视野常伴有镜面反光、部分遮挡和频繁组织接触。尽管近期机器人系统和基于视觉的辅助技术改善了操作员人机工程学并提供感知提示,但在显著解剖变异和安全关键视觉伪影下性能下降,阻碍了插管级操作的可靠自动化。本文提出BiliVLA,一种场景感知的视觉-语言-动作(VLA)框架,将胆道内镜导航建模为指令条件下的视觉运动学习问题。给定内镜观测和阶段特定语言指令,BiliVLA联合预测目标类别、具身边界框及连续内镜的离散三自由度(3-DoF)运动指令。该框架引入场景感知监督以提高语义目标一致性,并采用安全感知恢复监督,在管壁接触时诱导保守回撤行为。核心是两阶段训练范式:结合增强定位的监督微调(SFT)与组相对策略优化(GRPO),从而提升闭环导航中的动作可靠性与决策一致性。在三个ERCP子任务的物理仿真实验中,BiliVLA总体表现最优,总mIoU达0.9625,整体动作准确率为91.96%,整体成功率为84.85%。结果表明,融合语义定位、场景感知学习与奖励引导优化可强化感知-动作对齐,实现更鲁棒的自主胆道内镜导航。
原文摘要 · Abstract (English)
Endoscopic retrograde cholangiopancreatography (ERCP) demands precise endoscopic navigation and stable biliary cannulation within a narrow monocular field characterized by specular reflections, partial occlusions, and frequent tissue contact. Although recent robotic systems and vision-based assistance techniques improve operator ergonomics and provide perceptual cues, their performance degrades under pronounced anatomical variability and safety-critical visual artifacts, which hinders reliable autonomy in cannulation-grade procedures. Here, we present BiliVLA, a scene-aware Vision-Language-Action (VLA) framework that formulates biliary endoscopic navigation as an instruction-conditioned visuomotor learning problem. Given an endoscopic observation and a stage-specific language instruction, BiliVLA jointly predicts the target category, a grounded bounding box, and a discrete three-degree-of-freedom (3-DoF) motor command for a continuum endoscope. The proposed framework incorporates scene-aware supervision to improve semantic target consistency and safety-aware recovery supervision to induce conservative retreat behaviors under luminal wall contact. A key component of BiliVLA is a two-stage training paradigm that combines grounding-enhanced supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO), thereby improving action reliability and decision consistency during closed-loop navigation. Across three ERCP subtasks, BiliVLA achieves the best overall performance in physical phantom experiments, with a total mIoU of 0.9625, an overall action precision of 91.96\%, and an overall success rate (SR) of 84.85\%. These results indicate that integrating semantic grounding, scene-aware learning, and reward-guided optimization strengthens perception--action alignment and enables more robust autonomous biliary endoscopic navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。