用医学视觉语言模型提升机器人手术的决策能力
Medical Vision Language Models as Policies for Robotic Surgery
- 将医学视觉语言模型与强化学习结合,生成高层规划指令
- 在5个手术环境上成功率超70%,性能提升最高达11倍
- 适合需要医疗知识融合的智能手术系统研究者
基于视觉的近端策略优化(PPO)在视觉输入维度高、奖励稀疏且难以从原始视觉数据中提取任务相关特征的腹腔镜手术环境中表现受限。本文提出一种简单方法,将医学领域专用的视觉语言模型MedFlamingo与PPO结合。在LapGym中的五个多样化腹腔镜手术任务环境中,仅使用内窥镜视觉观测进行评估。MedFlamingo PPO优于标准视觉PPO和OpenFlamingo PPO基线,收敛更快,所有环境任务成功率均超过70%,相比基线提升幅度为66.67%至1114.29%。该方法通过每轮任务仅处理一次观察与指令,生成高层规划令牌,高效融合医学专业知识与实时视觉反馈。结果表明,专业化医疗知识对机器人手术规划与决策具有显著价值。
原文摘要 · Abstract (English)
Vision-based Proximal Policy Optimization (PPO) struggles with visual observation-based robotic laparoscopic surgical tasks due to the high-dimensional nature of visual input, the sparsity of rewards in surgical environments, and the difficulty of extracting task-relevant features from raw visual data. We introduce a simple approach integrating MedFlamingo, a medical domain-specific Vision-Language Model, with PPO. Our method is evaluated on five diverse laparoscopic surgery task environments in LapGym, using only endoscopic visual observations. MedFlamingo PPO outperforms and converges faster compared to both standard vision-based PPO and OpenFlamingo PPO baselines, achieving task success rates exceeding 70% across all environments, with improvements ranging from 66.67% to 1114.29% compared to baseline. By processing task observations and instructions once per episode to generate high-level planning tokens, our method efficiently combines medical expertise with real-time visual feedback. Our results highlight the value of specialized medical knowledge in robotic surgical planning and decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。