大模型在结直肠息肉光学诊断中表现接近专家,但仍有提升空间。
Performance of large language models in the optical diagnosis of colorectal polyps
- 用多模态大模型分析白光与窄带成像图片,自动分类息肉类型。
- Gemini 2.5 Pro 和 Claude Opus 4 在亚型区分上准确率最高,达41.7%正确率。
- 虽接近专家水平,但敏感性和特异性未达临床标准,需人机协作落地。
背景与研究目的:准确的结直肠息肉光学诊断可指导切除策略与随访方案,多模态大语言模型(MLLMs)在图像诊断中展现出潜力。本研究旨在评估MLLMs在分类结直肠息肉及预测组织学方面的诊断准确性。方法:基于PRIME数据集(包含白光与窄带成像图像),对Claude Opus 4、Google Gemini 2.5 Pro、GPT-o3、GPT-4o和GPT-5进行回顾性诊断性能评估。针对Paris、NICE分类及预测组织学,计算各模型的F1分数、正确率百分比及准确率,并与专家判断对比,共132例。采用Cochran's Q和McNemar检验分析模型间差异。结果:所有模型在良恶性分类上的F1分数均超过0.9;Gemini 2.5 Pro在侵袭性与非侵袭性息肉、低级别与高级别腺瘤区分上表现最优,分别为0.560和0.492。Claude Opus 4与GPT-5在Paris分类中的正确率显著高于其他模型,达41.7%。结论:Claude Opus 4与Gemini 2.5 Pro在区分息肉亚型方面表现最佳,最接近专家共识。然而其敏感性和特异性未达到ESGE标准,提示需开展前瞻性多中心试验,并设计人机协同工作流后方可临床部署。
原文摘要 · Abstract (English)
Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were >0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。