arXiv:2510.17038cs.ROcs.AI2025-10

用视觉和操作数据训练机器人导管自主导航,减少医生依赖。

DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation

  • 融合图像与手柄动作数据,构建多模态联合表征空间。
  • 在模拟血管模型上实现高精度动作预测,性能媲美纯动作基线。
  • 适合想开发自主医疗机器人的研究者或临床工程师参考。

心脏导管术是微创介入治疗的核心,但仍高度依赖人工操作。尽管机器人平台有所发展,现有系统多为“跟随-领导”模式,需持续医生输入,缺乏智能自主性,导致操作疲劳、辐射暴露增加及手术结果波动。本文提出DINO-CVA,一种多模态目标条件行为克隆框架,将视觉观测与手柄运动学信息融合至联合嵌入空间,使策略同时具备视觉感知与运动感知能力。动作通过专家示范自回归预测,目标条件引导导航至指定位置。搭建了基于合成血管模型的机器人实验平台,用于采集多模态数据并评估性能。结果表明,DINO-CVA在动作预测上达到高精度,性能与仅使用运动学的基线相当,且预测结果还基于解剖环境进行约束。这些发现证实了多模态、目标条件架构在导管导航中的可行性,是降低操作依赖、提升导管治疗可靠性的关键一步。

原文摘要 · Abstract (English)

Cardiac catheterization remains a cornerstone of minimally invasive interventions, yet it continues to rely heavily on manual operation. Despite advances in robotic platforms, existing systems are predominantly follow-leader in nature, requiring continuous physician input and lacking intelligent autonomy. This dependency contributes to operator fatigue, more radiation exposure, and variability in procedural outcomes. This work moves towards autonomous catheter navigation by introducing DINO-CVA, a multimodal goal-conditioned behavior cloning framework. The proposed model fuses visual observations and joystick kinematics into a joint embedding space, enabling policies that are both vision-aware and kinematic-aware. Actions are predicted autoregressively from expert demonstrations, with goal conditioning guiding navigation toward specified destinations. A robotic experimental setup with a synthetic vascular phantom was designed to collect multimodal datasets and evaluate performance. Results show that DINO-CVA achieves high accuracy in predicting actions, matching the performance of a kinematics-only baseline while additionally grounding predictions in the anatomical environment. These findings establish the feasibility of multimodal, goal-conditioned architectures for catheter navigation, representing an important step toward reducing operator dependency and improving the reliability of catheterbased therapies.

医疗机器人导管导航多模态学习行为克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。