用两阶段微调框架,让眼科影像模型更省资源、效果更好
TAPE: A two-stage parameter-efficient adaptation framework for foundation models in OCT-OCTA analysis
- 分两步适应:先对齐医学图像领域,再聚焦具体任务
- 仅用少量参数调整,就能在多种眼病上达到顶尖性能
- 适合算力有限但需高精度诊断的临床场景
光学相干断层扫描(OCT)和OCT血管成像(OCTA)的自动化分析对眼科诊断至关重要。现有主流方法从头训练依赖海量数据与大模型规模,限制了其在资源受限临床环境中的部署。尽管基于基础模型(FMs)的迁移学习有前景,但仍面临领域偏移与任务不匹配问题。为此,我们提出TAPE:一种基于参数高效微调的两阶段适配框架,将下游分割任务分解为领域对齐与任务拟合两个阶段。领域适配阶段首次在医学图像领域应用掩码图像建模下的参数高效微调(PEFT),为当前最佳实践。TAPE应用于通用(掩码自编码器,MAE)与专用(RETFound)基础模型,在视网膜层分割任务中展现出卓越的参数效率与跨病理泛化能力,达到当前最优性能。
原文摘要 · Abstract (English)
Automated analysis of optical coherence tomography (OCT) and OCT angiography (OCTA) images is critical for robust ophthalmic diagnosis. Existing mainstream methods trained from scratch rely heavily on massive data and model scale, thereby hindering their practical deployment in resource-constrained clinical settings. Although transfer learning based on foundation models (FMs) is promising, it still faces significant challenges: domain shift and task misalignment. To address these, we propose TAPE: A Two-stage Adaptation Framework via Parameter-Efficient Fine-tuning, which strategically decouples adaptation into domain alignment and task fitting for downstream segmentation. The domain adaptation stage notably applies parameter-efficient fine-tuning (PEFT) in the context of masked image modeling for medical image domain adaptation, a novel approach to the best of our knowledge. Applying TAPE to retinal layer segmentation on both universal (masked auto-encoder, MAE) and specialized (RETFound) FMs, it demonstrates superior parameter efficiency and achieves state-of-the-art generalization performance across diverse pathologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。