arXiv:2607.02692cs.CV2026-07

用视觉变压器和集成学习,自动融合眼底图与临床数据检测青光眼。

An Automated Multimodal Glaucoma Detection Framework Using ViT and a Stacking-Based Ensemble

  • 用ViT提取眼底图像特征,结合临床数据做多模态融合分类。
  • 样本级检测准确率97.47%,患者级检测达98.97%准确率与F1-score。
  • 适合医疗筛查系统开发,可部署为端到端辅助诊断平台。

青光眼是一种进行性眼病,若未能早期发现可能导致不可逆的视力丧失。传统诊断流程耗时且依赖专家判断,难以大规模推广。本研究在两种评估设置下开展青光眼检测:样本级(独立分析每个样本)和患者级(聚合每位患者的多份数据进行最终预测)。提出一种自动化多模态框架,整合眼底图像与临床数据。样本级检测中分别使用眼底图、临床特征及二者融合;患者级则聚合每位患者的多张眼底图表示与对应临床信息。采用视觉变压器(ViT)提取深层视觉特征,再由经典机器学习模型分类,并通过堆叠集成方法组合三个表现最佳的分类器以优化性能。在公开数据集PAPILA上的实验表明,样本级多模态分类达到97.47%准确率与97.50% F1-score,患者级检测准确率与F1-score均为98.97%。该框架已部署为端到端网络平台,支持自动化青光眼筛查与临床决策辅助。

原文摘要 · Abstract (English)

Glaucoma is a progressive eye disease that can lead to irreversible vision loss if not detected at an early stage. Conventional diagnostic procedures are often time-consuming and rely heavily on expert interpretation, limiting their scalability for large-scale screening. In this study, glaucoma detection is investigated under two evaluation settings: sample-wise, where individual samples are analyzed independently, and patient-wise, where data from each patient are aggregated for final prediction. An automated multimodal framework is proposed that integrates fundus images with clinical data. Under the sample-wise setting, detection is performed using fundus images and clinical features individually, as well as through their multimodal combination. Under the patient-wise setting, predictions are obtained by aggregating multiple fundus image representations with corresponding clinical information for each patient. Deep visual features are extracted using a Vision Transformer (ViT) architecture and classified using classical machine-learning models, with a stacking-based ensemble of the three best-performing classifiers employed to optimize performance. Experiments conducted on the publicly available PAPILA dataset demonstrate strong diagnostic performance, achieving 97.47% accuracy and a 97.50% F1-score for sample-wise multimodal classification, and 98.97% accuracy and F1-score for subject-wise detection. The proposed framework is further deployed as an end-to-end web-based platform to support automated glaucoma screening and clinical decision support.

青光眼检测多模态学习ViT医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。