arXiv:2505.15206cs.ROcs.AI2025-05被引 5

让内镜机器人自动跟踪病灶和切口标记,减少医生负担。

EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy

  • 分两阶段训练,结合监督与强化学习提升泛化能力。
  • 能零样本适应复杂场景,精准追踪病灶和圆形标记线。
  • 适合研发内镜手术机器人或做智能辅助系统的团队。

在内镜手术中,自动追踪异常区域并沿环形切口标记移动,可显著减轻内镜医师的认知负担。然而,传统基于模型的流程对各组件(如检测、运动规划)依赖人工调参,难以融入高级内镜意图,导致跨场景泛化能力差。视觉-语言-动作(VLA)模型通过端到端融合视觉感知、语言理解与运动规划,可在无需手动重校准的情况下语义适配术者指令,具备潜在优势。但将其应用于机器人内镜面临胃肠道复杂动态解剖环境的独特挑战。为此,本文提出专为连续体机器人设计的EndoVLA模型。给定内镜图像和术者追踪指令,该模型完成三项核心任务:(1) 腺瘤追踪,(2) 异常黏膜区域的勾画与跟随,(3) 在环形切割过程中沿圆形标记行进。为应对数据稀缺与领域偏移问题,我们提出双阶段策略:在自建的EndoVLA-Motion数据集上进行监督微调,并采用任务感知奖励进行强化学习微调。该方法显著提升了内镜追踪性能,实现了在多样化场景及复杂序列任务中的零样本泛化能力。

原文摘要 · Abstract (English)

In endoscopic procedures, autonomous tracking of abnormal regions and following circumferential cutting markers can significantly reduce the cognitive burden on endoscopists. However, conventional model-based pipelines are fragile for each component (e.g., detection, motion planning) requires manual tuning and struggles to incorporate high-level endoscopic intent, leading to poor generalization across diverse scenes. Vision-Language-Action (VLA) models, which integrate visual perception, language grounding, and motion planning within an end-to-end framework, offer a promising alternative by semantically adapting to surgeon prompts without manual recalibration. Despite their potential, applying VLA models to robotic endoscopy presents unique challenges due to the complex and dynamic anatomical environments of the gastrointestinal (GI) tract. To address this, we introduce EndoVLA, designed specifically for continuum robots in GI interventions. Given endoscopic images and surgeon-issued tracking prompts, EndoVLA performs three core tasks: (1) polyp tracking, (2) delineation and following of abnormal mucosal regions, and (3) adherence to circular markers during circumferential cutting. To tackle data scarcity and domain shifts, we propose a dual-phase strategy comprising supervised fine-tuning on our EndoVLA-Motion dataset and reinforcement fine-tuning with task-aware rewards. Our approach significantly improves tracking performance in endoscopy and enables zero-shot generalization in diverse scenes and complex sequential tasks.

内镜机器人视觉语言动作零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。