arXiv:2608.20478cs.RO2026-08

用语言指令消除内镜进退动作混淆,提升精准控制能力。

EndoLIFT: Language-Disambiguated Latent-Conditioned Rectified Flow for Bidirectional Endoscopic Control

论文配图:EndoLIFT: Language-Disambiguated Latent-Conditioned Rectified Flow for Bidirectional Endoscopic Control
图 1 · 摘自论文原文
  • 结合语言指令与潜变量的连续动作生成,解决进退方向歧义。
  • 方向判断准确率提升11.1个百分点,错误前进减少83%。
  • 适合需要高精度双向控制的医疗机器人场景。

常规胃肠道内镜检查本质上是双向的:器械需推进至目标解剖结构,随后回撤或反折进行观察,而外部提示可能要求提前反转。当指令阶段变化早于视觉场景时,几乎相同的观测可能需要相反的轴向动作。我们识别并形式化了这种双向内镜控制中的意图混淆问题,称为意图歧义。提出EndoLIFT(内镜语言指令流与轨迹潜变量模型),一种融合显式语言意图条件与潜变量驱动的修正流动作专家的视觉-语言-动作策略。该策略接收RGB图像、语言指令及前一动作状态;一个32维变分轨迹潜变量随机条件化连续动作块生成。相同观测下指令切换实验表明,语言可独立于潜变量选择轴向模式。相比无潜变量条件的对照模型,EndoLIFT导航方向准确率提升11.1个百分点,错误前进减少83%。架构控制的1比特模式标志参考显示其泛化能力弱,而EndoLIFT在44种未见语言变体中仍保持82.8%的意图遵循准确率。闭环评估中,相较于无变分轨迹潜变量的EndoLIFT,其在已见结肠模型和未见肺、胃模型上成功率均提升30个百分点,并成功完成10/10猪气管离体试验。结果表明,语言指令负责意图选择,潜变量则贡献于方向正确性与稳健回撤。

原文摘要 · Abstract (English)

Routine gastrointestinal endoscopy is intrinsically bidirectional: the instrument is advanced to reach target anatomy and later withdrawn or retroflexed for inspection, while an external cue may require earlier reversal. When the requested phase changes before the visual scene does, nearly identical observations can require opposite axial actions. We identify and formalize this ambiguity in bidirectional endoscopic control as intent aliasing. We propose EndoLIFT (Endoscopic Language-Instruction Flow with Trajectory Latents), a vision-language-action policy that combines explicit language-based intent conditioning with a latent-conditioned rectified-flow action expert. The policy receives RGB, a language instruction, and the previous-action state; a 32-D variational trajectory latent stochastically conditions continuous action-chunk generation. Controlled same-observation instruction swaps establish that language selects the axial mode, independently of whether the trajectory latent is present. Relative to the matched model without latent conditioning, EndoLIFT improves navigation-direction accuracy by 11.1 percentage points and reduces wrong-direction advance by 83\%. An architecture-controlled 1-bit mode-flag reference exhibits weaker canonical-anchor switching, while EndoLIFT retains 82.8\% intent-following accuracy across 44 held-out linguistic variants. In closed-loop evaluation, EndoLIFT improves overall success by 30 percentage points over EndoLIFT w/o VTL on both the seen colon phantom and the unseen lung and stomach phantoms, and completes 10/10 ex-vivo porcine-trachea trials. These results separate language-based intent selection from the trajectory latent's contribution to directional correctness and robust retraction.

内镜控制语言指令潜变量机器人手术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。