arXiv:2410.09750cs.CVcs.AI2024-10被引 17

用大模型理解手术场景,让AI能看懂并回答手术图像问题。

Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models

  • 将手术图像视频融入语言模型,构建专用多模态模型
  • 在手术问答数据集上表现优于已有方法
  • 适合医疗AI、智能手术辅助系统开发者使用

由大语言模型驱动的对话代理正在改变我们与视觉数据的交互方式。近期,大视觉语言模型(LVLM)在图像和视频领域得到广泛研究,但多数聚焦于通用场景。本文提出专为手术场景设计的LVLM——Surgical-LLaVA,将手术图像与视频的视觉表征融入语言特征空间,并在手术场景指令跟随数据上进行微调。实验表明,Surgical-LLaVA在手术上下文中展现出出色的多模态对话能力,甚至能在未见指令上表现出多模态行为。我们在手术场景的视觉问答数据集上进行了定量评估,结果优于先前工作,表明该模型在应对更复杂手术任务方面具有潜力。

原文摘要 · Abstract (English)

Conversation agents powered by large language models are revolutionizing the way we interact with visual data. Recently, large vision-language models (LVLMs) have been extensively studied for both images and videos. However, these studies typically focus on common scenarios. In this work, we introduce an LVLM specifically designed for surgical scenarios. We integrate visual representations of surgical images and videos into the language feature space. Consequently, we establish a LVLM model, Surgical-LLaVA, fine-tuned on instruction following data of surgical scenarios. Our experiments demonstrate that Surgical-LLaVA exhibits impressive multi-modal chat abilities in surgical contexts, occasionally displaying multi-modal behaviors on unseen instructions. We conduct a quantitative evaluation of visual question-answering datasets for surgical scenarios. The results show superior performance compared to previous works, indicating the potential of our model to tackle more complex surgery scenarios.

手术理解多模态视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。