arXiv:2505.12890cs.CV2025-05被引 4

打造专用于手术室的智能模型,让机器人更懂手术现场。

Specialized Foundation Models for Intelligent Operating Rooms

  • 融合视觉、声音和结构数据,构建手术室多模态理解模型。
  • 在手术场景下表现远超通用AI模型,识别更准更稳。
  • 提供多种尺寸模型,适配不同医院计算条件。

外科手术在复杂环境中进行,需协调手术团队、器械、影像及日益普及的智能机器人系统。未来手术室的安全与效率依赖于能理解复杂术中活动与风险的智能系统。现有计算方法缺乏全面性与泛化能力。本文提出ORQA,一种统一视觉、听觉与结构化数据的多模态基础模型,实现对术中场景的全景理解。其问答框架支持多样任务,可作为各类外科技术的智能核心。我们在多个临床场景下对比ORQA与通用视觉语言模型(如ChatGPT、Gemini),发现后者难以识别手术场景,而ORQA展现出显著更强且一致的表现。针对临床部署的多样性需求,我们设计并发布一系列轻量级ORQA模型,以适应不同算力环境。本工作为下一代智能手术解决方案奠定基础,助力医疗团队与设备厂商构建更智能、更安全的手术室。

原文摘要 · Abstract (English)

Surgical procedures unfold in complex environments demanding coordination between surgical teams, tools, imaging and increasingly, intelligent robotic systems. Ensuring safety and efficiency in ORs of the future requires intelligent systems, like surgical robots, smart instruments and digital copilots, capable of understanding complex activities and hazards of surgeries. Yet, existing computational approaches, lack the breadth, and generalization needed for comprehensive OR understanding. We introduce ORQA, a multimodal foundation model unifying visual, auditory, and structured data for holistic surgical understanding. ORQA's question-answering framework empowers diverse tasks, serving as an intelligence core for a broad spectrum of surgical technologies. We benchmark ORQA against generalist vision-language models, including ChatGPT and Gemini, and show that while they struggle to perceive surgical scenes, ORQA delivers substantially stronger, consistent performance. Recognizing the extensive range of deployment settings across clinical practice, we design, and release a family of smaller ORQA models tailored to different computational requirements. This work establishes a foundation for the next wave of intelligent surgical solutions, enabling surgical teams and medical technology providers to create smarter and safer operating rooms.

手术机器人多模态智能手术室

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。