用多视角眼底图像检测中风,提出首个视网膜Transformer模型。
Braided Vision Transformer for Stroke Detection in Multi-view Retinal Fundus Imaging

- 设计编织式视觉变换器,融合双眼前后视角特征
- 在自建数据集上实现0.75的AUC,优于常规ViT
- 适合眼科筛查与中风早期预警场景
中风仍是全球主要致死致残原因,亟需快速精准评估。眼底成像因其能反映脑血管和神经风险因素,成为中风评估的有前景手段。相较于传统神经影像,眼底成像具有无创、低成本、便携等优势,适用于快速筛查。本文探索了基于双眼黄斑中心与视乳头中心视角的眼底图像在中风及短暂性脑缺血发作(TIA)检测中的可行性。我们首次提出了用于眼底成像中风评估的视觉变换器模型——编织式视觉变换器(Braided Vision Transformer, BViT),通过同时提取多视角图像特征并捕捉双眼中跨视角关系,实现对与脑血管事件相关视网膜生物标志物更深入的理解。在自建的Stroke-Data数据集上的实验表明,BViT在中风检测任务中达到0.75的AUC,显著优于常规视觉变换器。
原文摘要 · Abstract (English)
Stroke remains a leading cause of mortality and morbidity worldwide, emphasizing the importance of its accurate and immediate assessment. Retinal fundus imaging has emerged as a promising modality for stroke assessment, as the retina reflects cerebrovascular and neurological risk factors. Contrary to conventional neuroimaging techniques, retinal fundus imaging offers a non-invasive, cost-effective, and portable alternative for rapid screening. This paper explores the feasibility of retinal fundus imaging for stroke and transient ischemic attack (TIA) detection using macula-centric and optic nerve head-centric views captured from both eyes. Our study introduces, to the best of our knowledge, the first vision transformer model for retinal fundus imaging in stroke assessment, offering a novel approach for capturing retinal patterns. Thereby, we propose the Braided Vision Transformer (BViT) model, which extracts representative features from the given multi-view images while simultaneously capturing inter-view relationships across both eyes, enabling a more informative understanding of retinal biomarkers associated with cerebrovascular events. Experiments conducted on our collected Stroke-Data dataset demonstrate that BViT achieves an AUC score of 0.75 for stroke detection, outperforming regular vision transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。