arXiv:2409.17332cs.CVcs.AI2024-09被引 3

用自监督学习和块扩展,让通用视觉模型适配眼底影像,不丢旧知识且更省数据。

Block Expanded DINORET: Adapting Natural Domain Foundation Models for Retinal Imaging Without Catastrophic Forgetting

  • 提出块扩展策略,提升自然域模型在眼底图像上的适应能力。
  • 新模型在糖尿病视网膜病变和青光眼检测中表现优于现有方法,且少样本下更高效。
  • 有效缓解微调中的灾难性遗忘,适合医疗单位定制化模型部署。

将深度学习融入医学影像有望大幅提升诊断水平,但泛化能力仍受限。基于自监督学习的基座模型可改善数据效率。自然域基座模型在医学影像中具潜力,但其领域适应(尤其结合自监督学习与参数高效微调)及微调时的灾难性遗忘问题研究尚不充分。本文基于DINOv2视觉变换器,通过自监督学习构建两个新型基座模型:DINORET与BE DINORET,用于眼底图像分类任务。使用公开的眼底彩色照片进行模型训练与微调,应用于糖尿病视网膜病变分期与青光眼检测。引入块扩展作为新颖的领域适应策略,并评估了灾难性遗忘情况。模型在多个数据集上表现优异,块扩展模型多数任务得分最高,且成功缓解灾难性遗忘。少样本学习实验表明,DINORET与BE DINORET在数据效率上超越RETFound。本研究证明,通过自监督学习与块扩展,可有效将自然域视觉模型迁移到眼底影像,实现性能稳健、不丢失原有能力,助力医疗机构为本地患者群体定制视觉模型,推动全球医疗普惠。

原文摘要 · Abstract (English)

Integrating deep learning into medical imaging is poised to greatly advance diagnostic methods but it faces challenges with generalizability. Foundation models, based on self-supervised learning, address these issues and improve data efficiency. Natural domain foundation models show promise for medical imaging, but systematic research evaluating domain adaptation, especially using self-supervised learning and parameter-efficient fine-tuning, remains underexplored. Additionally, little research addresses the issue of catastrophic forgetting during fine-tuning of foundation models. We adapted the DINOv2 vision transformer for retinal imaging classification tasks using self-supervised learning and generated two novel foundation models termed DINORET and BE DINORET. Publicly available color fundus photographs were employed for model development and subsequent fine-tuning for diabetic retinopathy staging and glaucoma detection. We introduced block expansion as a novel domain adaptation strategy and assessed the models for catastrophic forgetting. Models were benchmarked to RETFound, a state-of-the-art foundation model in ophthalmology. DINORET and BE DINORET demonstrated competitive performance on retinal imaging tasks, with the block expanded model achieving the highest scores on most datasets. Block expansion successfully mitigated catastrophic forgetting. Our few-shot learning studies indicated that DINORET and BE DINORET outperform RETFound in terms of data-efficiency. This study highlights the potential of adapting natural domain vision models to retinal imaging using self-supervised learning and block expansion. BE DINORET offers robust performance without sacrificing previously acquired capabilities. Our findings suggest that these methods could enable healthcare institutions to develop tailored vision models for their patient populations, enhancing global healthcare inclusivity.

眼底影像自监督学习模型迁移灾难性遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。