arXiv:2606.30698cs.RO2026-06中稿 · IEEE/RSJ IROS 2026

用视觉语言模型实现导丝导航的动态奖励调整,提升复杂血管手术的自主性。

Vision-Language Procedural Reasoning for Context-Aware Reward Modeling of Robotic Endovascular Guidewire Navigation

论文配图:Vision-Language Procedural Reasoning for Context-Aware Reward Modeling of Robotic Endovascular Guidewire Navigation
图 1 · 摘自论文原文
  • 引入多模态大模型解析实时图像,推理手术阶段与解剖上下文
  • 根据推理结果动态调整奖励权重,解决不同阶段目标冲突问题
  • 在真实机器人平台验证,显著提升导航可靠性和效率

机器人辅助血管介入手术需要在复杂且患者特异的血管解剖结构中实现精确、稳定且上下文感知的导丝导航。尽管机器人精度和基于学习的控制已有进展,现有自主导航方法仍受限于静态奖励函数,缺乏对解剖上下文与任务进展的显式过程推理。为此,本文提出一种视觉-语言过程推理(VL-PR)框架,集成多模态大语言模型(MLLM)作为过程推理模块,实时解析视觉观测以推断高层导航上下文。该模块不生成低层控制指令,而是通过动态调整各奖励成分的重要性,实现上下文感知的奖励自适应,使单一策略能应对竞争性目标与复杂状态转换,同时保持全局任务一致性。在多种血管场景的真实机器人平台上实验表明,该方法显著提升了任务可靠性与导航效率,优于静态奖励方法,为复杂多任务机器人血管手术提供了可扩展解决方案。

原文摘要 · Abstract (English)

Robotic-assisted endovascular interventions demand accurate, stable, and context-aware guidewire navigation in complex and patient-specific vascular anatomies. Despite recent advances in robotic precision and learning-based control, existing autonomous navigation methods remain limited by their reliance on static reward functions and the lack of explicit procedural reasoning regarding anatomical context and task progression. To address these challenges, this paper proposes a vision-language procedural reasoning (VL-PR) framework for autonomous guidewire navigation. The framework integrates a multimodal large language model (MLLM) as a procedural reasoning module that interprets real-time visual observations to infer high-level navigation contexts. Instead of generating low-level control commands, the inferred procedural insights enable context-aware reward adaptation by dynamically adjusting the importance of reward components across different navigation phases. This approach allows a single policy to resolve competing objectives and handle complex transitions while preserving a consistent global task goal. Experiments on a physical robotic platform across diverse vascular scenarios demonstrate enhanced task reliability and streamlined navigational efficiency, highlighting the advantages over static-reward methods and offering a scalable solution for complex and multi-task robotic endovascular procedures.

机器人手术视觉语言模型强化学习导丝导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。