融合CNN与Transformer,提升眼底病灶检测精度
Ophthalmic Biomarker Detection with Parallel Prediction of Transformer and Convolutional Architecture
- 并行使用CNN与Vision Transformer提取局部与全局特征
- 在OLIVES数据集上6类病灶检测的宏平均F1显著提升
- 适合医学图像分析、眼科AI研发人员参考
眼科疾病是全球重大健康问题,亟需精准诊断工具。光学相干断层扫描(OCT)提供高分辨率视网膜横截面图像,已成为眼科关键影像手段。传统上医生需手动识别疾病与生物标志物,近年来深度学习被广泛用于医疗诊断任务,实现快速精准分析。本文提出一种新方法,通过卷积神经网络(CNN)与视觉Transformer的集成模型进行眼科生物标志物检测。CNN擅长捕捉图像局部特征,而Transformer能有效提取全局上下文信息。两者并行结合,可兼顾局部细节与整体结构。该方法在OLIVES数据集上用于检测6种主要生物标志物,结果显示宏平均F1分数显著提升。
原文摘要 · Abstract (English)
Ophthalmic diseases represent a significant global health issue, necessitating the use of advanced precise diagnostic tools. Optical Coherence Tomography (OCT) imagery which offers high-resolution cross-sectional images of the retina has become a pivotal imaging modality in ophthalmology. Traditionally physicians have manually detected various diseases and biomarkers from such diagnostic imagery. In recent times, deep learning techniques have been extensively used for medical diagnostic tasks enabling fast and precise diagnosis. This paper presents a novel approach for ophthalmic biomarker detection using an ensemble of Convolutional Neural Network (CNN) and Vision Transformer. While CNNs are good for feature extraction within the local context of the image, transformers are known for their ability to extract features from the global context of the image. Using an ensemble of both techniques allows us to harness the best of both worlds. Our method has been implemented on the OLIVES dataset to detect 6 major biomarkers from the OCT images and shows significant improvement of the macro averaged F1 score on the dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。