评测DeepSeek在手术视觉语言任务中的表现,发现其尚不满足临床需求。
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
- 用手术数据集测试DeepSeek模型的多模态理解能力
- 在单句问答任务中表现优于通用模型,但整体仍不足
- 强调需手术专用数据微调才能用于真实医疗场景
DeepSeek模型在通用场景理解、问答和文本生成任务中表现出色,得益于高效的训练范式和强大的推理能力。本研究评估了DeepSeek模型在机器人手术场景下的对话能力,聚焦于单短语问答、视觉问答和详细描述等任务,其中单短语问答进一步细分为器械识别、动作理解与空间位置分析。通过公开数据集EndoVis18和CholecT50及其对应对话数据进行广泛评估,结果表明,相较于现有通用多模态大模型,DeepSeek-VL2在复杂手术场景理解任务中表现更优;尽管DeepSeek-V3为纯语言模型,直接输入图像标记后在单句问答任务上亦有较好表现。然而,总体而言,DeepSeek模型仍无法满足手术场景理解的临床要求:在通用提示下缺乏对全局手术概念的有效分析,难以提供详尽的场景洞察。基于此,我们认为DeepSeek模型未经手术特定数据微调,尚不具备应用于手术视觉语言任务的能力。
原文摘要 · Abstract (English)
The DeepSeek models have shown exceptional performance in general scene understanding, question-answering (QA), and text generation tasks, owing to their efficient training paradigm and strong reasoning capabilities. In this study, we investigate the dialogue capabilities of the DeepSeek model in robotic surgery scenarios, focusing on tasks such as Single Phrase QA, Visual QA, and Detailed Description. The Single Phrase QA tasks further include sub-tasks such as surgical instrument recognition, action understanding, and spatial position analysis. We conduct extensive evaluations using publicly available datasets, including EndoVis18 and CholecT50, along with their corresponding dialogue data. Our empirical study shows that, compared to existing general-purpose multimodal large language models, DeepSeek-VL2 performs better on complex understanding tasks in surgical scenes. Additionally, although DeepSeek-V3 is purely a language model, we find that when image tokens are directly inputted, the model demonstrates better performance on single-sentence QA tasks. However, overall, the DeepSeek models still fall short of meeting the clinical requirements for understanding surgical scenes. Under general prompts, DeepSeek models lack the ability to effectively analyze global surgical concepts and fail to provide detailed insights into surgical scenarios. Based on our observations, we argue that the DeepSeek models are not ready for vision-language tasks in surgical contexts without fine-tuning on surgery-specific datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。