arXiv:2608.00586cs.CV2026-08

对比不同预训练策略对眼底图像迁移能力的影响

Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging

论文配图:Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging
图 1 · 摘自论文原文
  • 用局部块的多实例学习框架评估视觉变换器在眼底图中的表示迁移
  • 自监督预训练模型表现优于MAE,DINOv3达0.863的加权卡帕值
  • 适合关注医学影像迁移学习与模型选择的研究者

尽管基础模型被广泛用作医学影像的特征提取器,但不同预训练策略如何影响其在弱监督眼科影像任务中的可迁移性仍不清楚。本研究在超广角眼底成像中,通过基于局部块的多实例学习(MIL)框架,评估了多种预训练视觉变换器的表示性能。比较了在ImageNet-1k上使用监督、掩码自编码器(MAE)和自蒸馏目标预训练的ViT-B编码器,保持下游聚合架构一致。结果显示,预训练目标显著影响冻结表示的迁移效果,监督和自蒸馏模型优于MAE。更大规模预训练的DINOv3模型表现最佳,五分类糖尿病视网膜病变分级的加权卡帕值达0.863,接近DINOv1。注意力分析揭示不同预训练表示具有不同的局部块聚合行为,部分微调可显著缩小MAE模型的性能差距。结果表明,预训练策略既影响表示迁移性,也影响MIL中局部证据的聚合方式,进而决定下游分类性能。

原文摘要 · Abstract (English)

Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pretraining strategies influence the transferability of learned representations to weakly supervised ophthalmic imaging tasks. We investigate this question in ultra-widefield (UWF) retinal imaging by evaluating foundation model representations within a patch-based multiple instance learning (MIL) framework for disease classification on UWF images. We compare Vision Transformer encoders pretrained with supervised, Masked Autoencoder (MAE), and self-distillation objectives, while keeping the downstream aggregation architecture unchanged. Within a controlled comparison of ViT-B encoders pretrained on ImageNet-1k, the choice of pretraining objective substantially influenced frozen representation transfer, with supervised and self-distillation-based models outperforming MAE. A contemporary DINOv3 model pretrained at a larger scale achieved the strongest overall performance, with a quadratic weighted kappa of 0.863 for five-class diabetic retinopathy grading, comparable with DINOv1. Attention analysis further revealed distinct patch-aggregation behaviours associated with the different pretrained representations, while partial fine-tuning substantially reduced the performance gap for MAE. These findings suggest that pretraining strategy influences both representation transferability and the subsequent aggregation of patch-level evidence within MIL, resulting in differences in downstream classification performance.

眼底成像迁移学习视觉变换器自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。