用轻量LoRA与联邦学习实现隐私保护的视网膜病诊断,兼顾准确与可解释。
Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis
- 融合多尺度嵌入与LoRA,提升特征捕捉效率
- 在OCTDL和眼病数据集上AUC超主流模型
- 支持分布式训练,适合医疗数据隐私场景
准确且隐私安全的眼科疾病诊断仍是医学影像中的关键挑战,尤其受限于现有深度学习模型在数据不平衡、数据隐私、空间特征多样性及临床可解释性方面的不足。本文提出一种基于DeiT的新型数据高效框架,集成上下文感知的多尺度图像块嵌入、低秩适配(LoRA)、知识蒸馏与联邦学习,统一应对上述问题。该模型通过多尺度图像块表示与局部-全局注意力机制,有效捕捉视网膜的局部与全局特征。引入LoRA显著减少可训练参数,提升计算效率;联邦学习实现无数据泄露的分布式训练,保障隐私安全。知识蒸馏策略进一步增强数据稀缺场景下的泛化能力。在两个基准数据集OCTDL和眼病图像数据集上的全面评估显示,该框架在AUC、F1分数和精确率等关键指标上持续优于传统CNN与先进Transformer架构。此外,Grad-CAM++可视化提供可解释的预测依据,增强临床可信度。本工作为眼科诊断中可扩展、安全、可解释的AI应用奠定坚实基础。
原文摘要 · Abstract (English)
Accurate and privacy-preserving diagnosis of ophthalmic diseases remains a critical challenge in medical imaging, particularly given the limitations of existing deep learning models in handling data imbalance, data privacy concerns, spatial feature diversity, and clinical interpretability. This paper proposes a novel Data efficient Image Transformer (DeiT) based framework that integrates context aware multiscale patch embedding, Low-Rank Adaptation (LoRA), knowledge distillation, and federated learning to address these challenges in a unified manner. The proposed model effectively captures both local and global retinal features by leveraging multi scale patch representations with local and global attention mechanisms. LoRA integration enhances computational efficiency by reducing the number of trainable parameters, while federated learning ensures secure, decentralized training without compromising data privacy. A knowledge distillation strategy further improves generalization in data scarce settings. Comprehensive evaluations on two benchmark datasets OCTDL and the Eye Disease Image Dataset demonstrate that the proposed framework consistently outperforms both traditional CNNs and state of the art transformer architectures across key metrics including AUC, F1 score, and precision. Furthermore, Grad-CAM++ visualizations provide interpretable insights into model predictions, supporting clinical trust. This work establishes a strong foundation for scalable, secure, and explainable AI applications in ophthalmic diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。