用自监督学习融合多源OCT数据,提升小样本下眼病分类性能
Multi-OCT-SelfNet: Integrating Self-Supervised Learning with Multi-Source Data Fusion for Enhanced Multi-Class Retinal Disease Classification
- 通过自监督预训练+微调两阶段框架,融合多源OCT数据增强表征能力
- 在低数据量和跨域场景下仍保持稳定性能,优于ResNet-50基线
- 适合数据稀缺或临床部署环境变化大的眼病智能诊断应用
医疗领域因隐私问题难以获取大规模数据,但深度学习模型的训练又依赖大量数据。为应对数据匮乏挑战,本文结合多种数据源,利用基于大语言模型的SwinV2自监督框架,从多模态OCT数据中学习更具泛化能力的表征表示,提升模型对新数据的外推能力。采用两阶段训练策略:自监督预训练与下游监督微调。在三个数据集上进行消融实验,对比不同编码器、无数据融合、低数据可用性及无自监督预训练等场景,结果表明本方法在各类条件下均表现稳健,显著优于基准模型ResNet-50。
原文摘要 · Abstract (English)
In the medical domain, acquiring large datasets poses significant challenges due to privacy concerns. Nonetheless, the development of a robust deep-learning model for retinal disease diagnosis necessitates a substantial dataset for training. The capacity to generalize effectively on smaller datasets remains a persistent challenge. The scarcity of data presents a significant barrier to the practical implementation of scalable medical AI solutions. To address this issue, we've combined a wide range of data sources to improve performance and generalization to new data by giving it a deeper understanding of the data representation from multi-modal datasets and developed a self-supervised framework based on large language models (LLMs), SwinV2 to gain a deeper understanding of multi-modal dataset representations, enhancing the model's ability to extrapolate to new data for the detection of eye diseases using optical coherence tomography (OCT) images. We adopt a two-phase training methodology, self-supervised pre-training, and fine-tuning on a downstream supervised classifier. An ablation study conducted across three datasets employing various encoder backbones, without data fusion, with low data availability setting, and without self-supervised pre-training scenarios, highlights the robustness of our method. Our findings demonstrate consistent performance across these diverse conditions, showcasing superior generalization capabilities compared to the baseline model, ResNet-50.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。