arXiv:2409.19590cs.RO2024-09被引 42

用视觉语言动作模型实现手术器械精准递送,支持语音指令实时操作。

RoboNurse-VLA: Robotic Scrub Nurse System based on Vision-Language-Action Model

论文配图:RoboNurse-VLA: Robotic Scrub Nurse System based on Vision-Language-Action Model
图 1 · 摘自论文原文
  • 基于视觉-语言-动作模型,结合SAM 2与Llama 2实现多模态理解
  • 在未见器械和复杂场景下仍保持高成功率的器械递送能力
  • 适用于需要高精度、实时响应的智能手术辅助场景

现代医疗中,对自主机器人助手的需求日益增长,尤其是在要求高精度和可靠性的手术室中。机器人洗手护士已成为提升手术效率、减少人为错误的有前景解决方案。然而,在动态环境中准确抓取和递送手术器械,特别是复杂或难以抓握的物品时仍面临挑战。本文提出一种基于视觉-语言-动作(VLA)模型的新型机器人洗手护士系统RoboNurse-VLA,集成分割一切模型2(SAM 2)与Llama 2语言模型。该系统可根据外科医生的语音指令,实现实时、高精度的手术器械抓取与递送。利用先进的视觉与语言模型,系统有效解决了物体检测、位姿优化及复杂器械处理等关键问题。通过大量评估,RoboNurse-VLA在未见过的器械和挑战性物品上均表现出色,显著优于现有模型,实现了高成功率的器械交接。本工作为自主手术辅助迈出重要一步,展示了VLA模型在真实医疗应用中的潜力。

原文摘要 · Abstract (English)

In modern healthcare, the demand for autonomous robotic assistants has grown significantly, particularly in the operating room, where surgical tasks require precision and reliability. Robotic scrub nurses have emerged as a promising solution to improve efficiency and reduce human error during surgery. However, challenges remain in terms of accurately grasping and handing over surgical instruments, especially when dealing with complex or difficult objects in dynamic environments. In this work, we introduce a novel robotic scrub nurse system, RoboNurse-VLA, built on a Vision-Language-Action (VLA) model by integrating the Segment Anything Model 2 (SAM 2) and the Llama 2 language model. The proposed RoboNurse-VLA system enables highly precise grasping and handover of surgical instruments in real-time based on voice commands from the surgeon. Leveraging state-of-the-art vision and language models, the system can address key challenges for object detection, pose optimization, and the handling of complex and difficult-to-grasp instruments. Through extensive evaluations, RoboNurse-VLA demonstrates superior performance compared to existing models, achieving high success rates in surgical instrument handovers, even with unseen tools and challenging items. This work presents a significant step forward in autonomous surgical assistance, showcasing the potential of integrating VLA models for real-world medical applications. More details can be found at https://robonurse-vla.github.io.

手术机器人多模态感知语音控制智能辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。