arXiv:2502.12379cs.CVcs.LG2025-02被引 1

用自研OCT数据训练ViT,发现数据够多时无需ImageNet预训练

OCT Data is All You Need: How Vision Transformers with and without Pre-training Benefit Imaging

  • 直接在OCT数据上训练ViT,无需ImageNet预训练
  • 小样本时预训练加速收敛,大样本下精度相当甚至更优
  • 提示应构建OCT专用预训练模型,适合眼科影像研究者

光学相干断层扫描(OCT)提供高分辨率横截面图像,对多种疾病诊断有重要价值。然而,其图像特性与自然图像差异显著,引发关于基于ImageNet的预训练是否始终有益的疑问。本文研究了在不同数据规模下,ImageNet预训练对视觉变换器(ViT)在OCT图像分类任务上的影响。实验涵盖四种视网膜病理类型(CNV、DME、Drusen、Normal)。结果表明:尽管预训练在小样本时能加速收敛并可能提升性能,但在拥有足够OCT数据的情况下,从头训练可达到相当或更优的准确率。研究强调预训练需匹配领域特征,呼吁开展大规模OCT特定预训练的进一步探索。

原文摘要 · Abstract (English)

Optical Coherence Tomography (OCT) provides high-resolution cross-sectional images useful for diagnosing various diseases, but their distinct characteristics from natural images raise questions about whether large-scale pre-training on datasets like ImageNet is always beneficial. In this paper, we investigate the impact of ImageNet-based pre-training on Vision Transformer (ViT) performance for OCT image classification across different dataset sizes. Our experiments cover four-category retinal pathologies (CNV, DME, Drusen, Normal). Results suggest that while pre-training can accelerate convergence and potentially offer better performance in smaller datasets, training from scratch may achieve comparable or even superior accuracy when sufficient OCT data is available. Our findings highlight the importance of matching domain characteristics in pre-training and call for further study on large-scale OCT-specific pre-training.

OCT影像视觉变换器预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。