arXiv:2507.14738cs.CV2025-07

融合眼底影像与社会健康数据,提升糖尿病视网膜病变分期准确率。

MultiRetNet: A Multimodal Vision Model and Deferral System for Staging Diabetic Retinopathy

  • 整合眼底图像、社会经济因素与共病信息进行多模态融合分析。
  • 在低质量图像上保持诊断准确率,识别需医生复核的异常样本。
  • 适合关注医疗公平与早期筛查的临床研究者使用。

糖尿病视网膜病变(DR)是全球超1亿人面临可预防失明的主要原因。在美国,低收入群体因筛查受限,更易在诊断前进展至严重阶段,共病情况也加速疾病发展。本文提出MultiRetNet,一种结合眼底影像、社会经济因素及共病特征的多模态模型,并集成临床延后系统实现人机协同。实验比较三种多模态融合方法,发现全连接层融合最为稳健。通过生成对抗性与低质量图像并采用对比学习训练延后系统,使模型能识别分布外样本,提示需人工审查。该系统在劣质图像上仍保持诊断精度,整合关键健康数据,有助于改善弱势群体的早期检测,降低医疗成本,缓解医疗不平等,推动医疗公平。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is a leading cause of preventable blindness, affecting over 100 million people worldwide. In the United States, individuals from lower-income communities face a higher risk of progressing to advanced stages before diagnosis, largely due to limited access to screening. Comorbid conditions further accelerate disease progression. We propose MultiRetNet, a novel pipeline combining retinal imaging, socioeconomic factors, and comorbidity profiles to improve DR staging accuracy, integrated with a clinical deferral system for a clinical human-in-the-loop implementation. We experiment with three multimodal fusion methods and identify fusion through a fully connected layer as the most versatile methodology. We synthesize adversarial, low-quality images and use contrastive learning to train the deferral system, guiding the model to identify out-of-distribution samples that warrant clinician review. By maintaining diagnostic accuracy on suboptimal images and integrating critical health data, our system can improve early detection, particularly in underserved populations where advanced DR is often first identified. This approach may reduce healthcare costs, increase early detection rates, and address disparities in access to care, promoting healthcare equity.

糖尿病视网膜病变多模态融合医疗公平人工智能辅助诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。