arXiv:2507.01278cs.CL2025-07被引 1

GPT-4可基于眼底照片描述模拟眼科诊断,但精度有限。

Evaluating Large Language Models for Multimodal Simulated Ophthalmic Decision-Making in Diabetic Retinopathy and Glaucoma Screening

  • 用结构化文本描述眼底图像,让GPT-4判断糖尿病视网膜病变和青光眼风险。
  • 识别正常病例准确率达67.5%,青光眼判断准确率仅约78%且指标极低。
  • 真实或合成的临床数据对结果影响不大,适合用于教学或标注辅助。

大型语言模型(LLMs)可基于自然语言提示模拟临床推理,但在眼科领域的应用尚未深入探索。本研究评估了GPT-4在糖尿病视网膜病变(DR)和青光眼筛查中,基于眼底照相的结构化文本描述进行临床决策的能力,包括加入真实或合成临床元数据的影响。研究使用300张标注的眼底图像开展回顾性诊断验证。GPT-4接收包含图像描述及患者元数据的结构化提示,任务为分配ICDR严重程度评分、推荐是否转诊DR,以及估计青光眼筛查所需的杯盘比。性能通过准确率、宏/加权F1分数及Cohen's kappa评估。采用McNemar检验与变化率分析评估元数据影响。结果显示,GPT-4在ICDR分类中表现中等(准确率67.5%,宏F1 0.33,加权F1 0.67,kappa 0.25),主要得益于正常病例的正确识别;在二分类转诊任务中表现提升(准确率82.3%,F1 0.54,kappa 0.44)。青光眼转诊任务整体表现差(准确率约78%,F1<0.04,kappa<0.03)。元数据的加入未显著改变结果(McNemar p > 0.05),预测一致性高。结论:GPT-4可模拟基础眼科决策,但复杂任务精度不足。虽不适用于临床,但可能在教育、文档生成或图像标注流程中提供辅助。

原文摘要 · Abstract (English)

Large language models (LLMs) can simulate clinical reasoning based on natural language prompts, but their utility in ophthalmology is largely unexplored. This study evaluated GPT-4's ability to interpret structured textual descriptions of retinal fundus photographs and simulate clinical decisions for diabetic retinopathy (DR) and glaucoma screening, including the impact of adding real or synthetic clinical metadata. We conducted a retrospective diagnostic validation study using 300 annotated fundus images. GPT-4 received structured prompts describing each image, with or without patient metadata. The model was tasked with assigning an ICDR severity score, recommending DR referral, and estimating the cup-to-disc ratio for glaucoma referral. Performance was evaluated using accuracy, macro and weighted F1 scores, and Cohen's kappa. McNemar's test and change rate analysis were used to assess the influence of metadata. GPT-4 showed moderate performance for ICDR classification (accuracy 67.5%, macro F1 0.33, weighted F1 0.67, kappa 0.25), driven mainly by correct identification of normal cases. Performance improved in the binary DR referral task (accuracy 82.3%, F1 0.54, kappa 0.44). For glaucoma referral, performance was poor across all settings (accuracy ~78%, F1 <0.04, kappa <0.03). Metadata inclusion did not significantly affect outcomes (McNemar p > 0.05), and predictions remained consistent across conditions. GPT-4 can simulate basic ophthalmic decision-making from structured prompts but lacks precision for complex tasks. While not suitable for clinical use, LLMs may assist in education, documentation, or image annotation workflows in ophthalmology.

眼科诊断大模型应用医学决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。