arXiv:2510.00882cs.CVcs.LG2025-10中稿 · publication at the…被引 1

用3D注意力网络精准识别青光眼,兼顾解剖结构与诊断效率。

AI-CNet3D: An Anatomically-Informed Cross-Attention Network with Multi-Task Consistency Fine-tuning for 3D Glaucoma Classification

  • 融合3D CNN与交叉注意力,捕捉视网膜上下半区及视盘的特征
  • 通过一致性微调提升模型准确率,在双数据集上超越主流方法
  • 参数量减少100倍,适合临床部署且结果可解释性强

青光眼是一种进行性眼病,可导致视神经损伤和不可逆的视力丧失。光学相干断层扫描(OCT)提供高分辨率的3D视网膜与视神经图像,已成为诊断关键工具。然而,传统将3D OCT信息压缩为2D报告的做法常丢失关键结构细节。为此,我们提出一种新型混合深度学习模型——AI-CNet3D,将交叉注意力机制嵌入3D卷积神经网络,从视网膜上/下半区、视盘(ONH)和黄斑区域提取关键特征。引入通道注意力表示(CAREs)可视化注意力输出,并基于梯度加权类激活图(Grad-CAM)进行一致性多任务微调,提升性能、可解释性与解剖一致性。模型通过沿两轴分割体积并应用交叉注意力,增强对半区不对称性的捕捉能力。在两个大型数据集上的验证表明,其在所有关键指标上均优于现有注意力与卷积模型。此外,该模型计算高效,参数量相比其他注意力机制减少一百倍,同时保持高诊断性能与相近的GFLOPS。

原文摘要 · Abstract (English)

Glaucoma is a progressive eye disease that leads to optic nerve damage, causing irreversible vision loss if left untreated. Optical coherence tomography (OCT) has become a crucial tool for glaucoma diagnosis, offering high-resolution 3D scans of the retina and optic nerve. However, the conventional practice of condensing information from 3D OCT volumes into 2D reports often results in the loss of key structural details. To address this, we propose a novel hybrid deep learning model that integrates cross-attention mechanisms into a 3D convolutional neural network (CNN), enabling the extraction of critical features from the superior and inferior hemiretinas, as well as from the optic nerve head (ONH) and macula, within OCT volumes. We introduce Channel Attention REpresentations (CAREs) to visualize cross-attention outputs and leverage them for consistency-based multi-task fine-tuning, aligning them with Gradient-Weighted Class Activation Maps (Grad-CAMs) from the CNN's final convolutional layer to enhance performance, interpretability, and anatomical coherence. We have named this model AI-CNet3D (AI-`See'-Net3D) to reflect its design as an Anatomically-Informed Cross-attention Network operating on 3D data. By dividing the volume along two axes and applying cross-attention, our model enhances glaucoma classification by capturing asymmetries between the hemiretinal regions while integrating information from the optic nerve head and macula. We validate our approach on two large datasets, showing that it outperforms state-of-the-art attention and convolutional models across all key metrics. Finally, our model is computationally efficient, reducing the parameter count by one-hundred--fold compared to other attention mechanisms while maintaining high diagnostic performance and comparable GFLOPS.

青光眼分类3D医学影像交叉注意力可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。