arXiv:2505.17808cs.CVcs.AI2025-05被引 4

融合CNN与ViT的注意力模型,提升青光眼早期筛查准确率

An Attention Infused Deep Learning System with Grad-CAM Visualization for Early Screening of Glaucoma

  • 采用交叉注意力融合自定义CNN与Vision Transformer
  • 在ACRIMA和Drishti数据集上性能优于单一CNN或ViT模型
  • 通过Grad-CAM可视化关键病灶区域,辅助医生诊断

本研究将定制化的卷积神经网络与颠覆性的Vision Transformer相结合,并引入创新的交叉注意力模块。利用两个高性能人工智能数据集ACRIMA和Drishti进行训练。交叉注意力机制通过双向特征交换,使模型能够学习视网膜图像中临床相关区域。实验表明,该融合模型在青光眼检测任务中显著优于独立的基准CNN和ViT模型。

原文摘要 · Abstract (English)

This research work reveals the strengths of intertwining a deep custom convolutional neural network with a disruptive Vision Transformer, both fused together with a radical Cross-Attention module. Here, two high-yielding datasets for artificial intelligence models in detecting glaucoma, namely ACRIMA and Drishti, are utilized. The Cross-Attention mechanism facilitates the model in learning regions in the fundus that are clinically relevant through bidirectional feature exchange between CNN and ViT streams. Experiments clearly depict improved performance when compared to standalone baseline CNN and ViT models.

青光眼筛查视觉变压器注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。