用大模型自动分析眼底OCT图像,判断质量、诊断青光眼并生成结构化报告。
Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model
- 基于视觉语言模型,结合OCT图像与临床报告进行微调。
- 青光眼检测准确率达86%,图像质量识别特异性强(98%)。
- 可生成包含七个象限的视网膜神经纤维层变化报告,适合临床辅助诊断。
目的:开发一种可解释的多模态大语言模型(MM-LLM),用于筛查视盘OCT环形扫描的质量,并生成包含青光眼诊断及各象限视网膜神经纤维层(RNFL)变薄评估的结构化临床报告。方法:回顾性队列研究纳入1,310名受试者,共43,849张Spectralis视盘OCT环形扫描图像(1,331只青光眼眼,867只健康眼),数据来自DIGS和ADAGES队列。使用Llama 3.2 Vision-Instruct模型进行微调,训练数据包括配对的OCT图像与自动生成的结构化临床报告,涵盖全局及七象限RNFL变薄描述;低质量图像标记为不可用,并配以固定拒绝语句。在独立测试集上评估模型在图像质量评估、青光眼检测和七象限RNFL变薄分类三项任务的表现,指标包括准确率、敏感性、特异性、精确率和F1分数。文本生成质量通过标准文本评价指标评估。结果:模型在质量分诊任务中准确率为0.90,特异性达0.98;青光眼检测准确率为0.86(敏感性0.91,特异性0.73,F1分数0.91);RNFL变薄预测准确率介于0.83至0.94之间,全局及颞侧表现最优。文本生成指标显示与参考报告高度一致(BLEU: 0.82;ROUGE-1: 0.94;ROUGE-2: 0.87;ROUGE-L: 0.92;BERTScore-F1: 0.99)。结论:微调后的MM-LLM能基于OCT影像生成准确的临床描述,在识别图像质量问题和青光眼方面表现优异,并提供七象限的RNFL变薄描述,有助于支持临床OCT评估。
原文摘要 · Abstract (English)
Objective: To develop an explainable multimodal large language model (MM-LLM) that (1) screens optic nerve head (ONH) OCT circle scans for quality and (2) generates structured clinical reports that include glaucoma diagnosis and sector-wise retinal nerve fiber layer (RNFL) thinning assessments. Design: Retrospective cohort study of 1,310 subjects contributing 43,849 Spectralis ONH OCT circle scans (1,331 glaucomatous and 867 healthy eyes) from the DIGS and ADAGES cohorts. Methods: A MM-LLM (Llama 3.2 Vision-Instruct model) was fine-tuned to generate clinical descriptions of OCT imaging data. Training data included paired OCT images and automatically generated, structured clinical reports that described global and sectoral RNFL thinning. Poor-quality scans were labeled as unusable and paired with a fixed refusal statement. The model was evaluated on a held-out test set for three tasks: quality assessment, glaucoma detection, and RNFL thinning classification across seven anatomical sectors. Evaluation metrics included accuracy, sensitivity, specificity, precision, and F1-score. Model description quality was also evaluated using standard text evaluation metrics. Results: The model achieved 0.90 accuracy and 0.98 specificity for quality triage. For glaucoma detection, accuracy was 0.86 (sensitivity 0.91, specificity 0.73, F1-score 0.91). RNFL thinning prediction accuracy ranged from 0.83 to 0.94, with highest performance in global and temporal sectors. Text generation scores showed strong alignment with reference reports (BLEU: 0.82; ROUGE-1: 0.94; ROUGE-2: 0.87; ROUGE-L: 0.92; BERTScore-F1: 0.99). Conclusions: The fine-tuned MM-LLM generated accurate clinical descriptions based on OCT imaging. The model achieved high accuracy in identifying image quality issues and detecting glaucoma. The model also provided sectoral descriptions of RNFL thinning to help support clinical OCT evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。