专为内镜手术设计的多模态大模型,提升手术理解与人机协作能力。
EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery
- 构建专用数据集Surg-396K,通过结构化标注增强手术场景理解
- 提出多尺度视觉交互与视觉对比推理机制,提升模型表征与推理能力
- 在5种对话模式8项任务中达顶尖水平,获外科医生高度认可
近年来,多模态大语言模型(MLLMs)在辅助诊断与决策方面展现出巨大潜力。在机器人辅助手术中,MLLMs可作为有效的手术训练与指导工具。然而,当前仍缺乏针对临床应用中手术场景理解的专用MLLM。本文提出EndoChat,以应对外科医生在手术场景理解中遇到的多种对话范式与子任务。为训练该模型,我们通过新构建的流水线,从大规模内镜手术数据集中系统提取手术信息,并生成结构化标注,构建了Surg-396K数据集。此外,引入多尺度视觉标记交互机制与基于视觉对比的推理机制,增强模型的表征学习与推理能力。模型在五个对话范式与八项手术场景理解任务中均达到当前最优表现。同时,我们邀请专业外科医生进行评估,多数给予积极反馈,认为与EndoChat协作体验良好。总体表明,EndoChat在推动机器人辅助手术的培训与自动化方面具有巨大潜力。
原文摘要 · Abstract (English)
Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools for surgical training and guidance. However, there is still a lack of MLLMs specialized for surgical scene understanding in clinical applications. In this work, we introduce EndoChat to address various dialogue paradigms and subtasks in surgical scene understanding that surgeons encounter. To train our EndoChat, we construct the Surg-396K dataset through a novel pipeline that systematically extracts surgical information and generates structured annotations based on collected large-scale endoscopic surgery datasets. Furthermore, we introduce a multi-scale visual token interaction mechanism and a visual contrast-based reasoning mechanism to enhance the model's representation learning and reasoning capabilities. Our model achieves state-of-the-art performance across five dialogue paradigms and eight surgical scene understanding tasks. Additionally, we conduct evaluations with professional surgeons, most of whom provide positive feedback on collaborating with EndoChat. Overall, these results demonstrate that our EndoChat has great potential to significantly advance training and automation in robotic-assisted surgery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。