arXiv:2604.25720cs.CVcs.CL2026-04

用对话式AI分析眼底照片,自动诊断老年黄斑变性并解释原因。

Toward Multimodal Conversational AI for Age-Related Macular Degeneration

  • 基于视觉问答训练,让AI像医生一样与患者对话。
  • 在三个诊断任务中准确率最高,超越现有模型,尤其对色素异常识别出色。
  • 医生评分显示其解释更清晰可信,适合临床辅助诊断使用。

尽管深度学习在视网膜疾病检测中表现优异,但多数系统仅输出静态判断,缺乏临床推理与交互解释。本研究提出OcularChat,一个基于Qwen2.5-VL微调的多模态大语言模型,通过模拟患者-医生对话,利用彩色眼底照片(CFPs)进行年龄相关性黄斑变性(AMD)的视觉问答诊断。共生成705,850组模拟对话与46,167张CFPs用于训练,使OcularChat能识别关键AMD特征并生成合理诊断。在AREDS数据集上,其在高级AMD、色素异常和脂质沉积物大小三项任务中的准确率分别为0.954、0.849和0.678,显著优于现有多模态大模型;在AREDS2数据集上同样保持领先。三位独立眼科医师评估显示,OcularChat在高级AMD(3.503 vs. 2.833)、色素异常(3.272 vs. 2.828)、脂质沉积物大小(3.064 vs. 2.433)及整体印象(2.978 vs. 2.464)上的平均分均高于强基线模型(5分制)。除客观分类性能优异外,该模型还具备诊断推理、临床相关解释与互动对话能力,且主观评价得分高。结果表明,多模态大语言模型可实现准确、可解释且临床可用的AMD图像诊断与分类。

原文摘要 · Abstract (English)

Despite strong performance of deep learning models in retinal disease detection, most systems produce static predictions without clinical reasoning or interactive explanation. Recent advances in multimodal large language models (MLLMs) integrate diagnostic predictions with clinically meaningful dialogue to support clinical decision-making and patient counseling. In this study, OcularChat, an MLLM, was fine-tuned from Qwen2.5-VL using simulated patient-physician dialogues to diagnose age-related macular degeneration (AMD) through visual question answering on color fundus photographs (CFPs). A total of 705,850 simulated dialogues paired with 46,167 CFPs were generated to train OcularChat to identify key AMD features and produce reasoned predictions. OcularChat demonstrated strong classification performance in AREDS, achieving accuracies of 0.954, 0.849, and 0.678 for the three diagnostic tasks: advanced AMD, pigmentary abnormalities, and drusen size, significantly outperforming existing MLLMs. On AREDS2, OcularChat remained the top-performing method on all tasks. Across three independent ophthalmologist graders, OcularChat achieved higher mean scores than a strong baseline model for advanced AMD (3.503 vs. 2.833), pigmentary abnormalities (3.272 vs. 2.828), drusen size (3.064 vs. 2.433), and overall impression (2.978 vs. 2.464) on a 5-point clinical grading rubric. Beyond strong objective performance in AMD severity classification, OcularChat demonstrated the ability to provide diagnostic reasoning, clinically relevant explanations, and interactive dialogue, with high performance in subjective ophthalmologist evaluation. These findings suggest that MLLMs may enable accurate, interpretable, and clinically useful image-based diagnosis and classification of AMD.

医学AI多模态对话系统眼底病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。