arXiv:2512.01273cs.CV2025-12被引 1

融合卷积与注意力机制,提升眼底图像分析精度与效率

nnMobileNet++: Towards Efficient Hybrid Networks for Retinal Image Analysis

  • 用动态蛇形卷积捕捉病变边界,结合分阶段变压器建模全局上下文
  • 在多个眼底数据集上达到顶尖准确率,计算开销仍很低
  • 适合医疗影像轻量化模型开发与临床辅助诊断系统部署

眼底成像是早期发现眼部及系统性疾病的关键无创手段。深度学习,尤其是卷积神经网络(CNN),在自动眼底图像分析中取得显著进展,支持视网膜图像分类、病灶检测和血管分割等任务。作为轻量级网络的代表,nnMobileNet在多个眼底基准测试中表现出色,同时保持低计算成本。然而,纯卷积架构难以捕捉长程依赖关系,也难以建模不规则病灶和细长血管结构,而这些特征对可靠临床诊断至关重要。为此,我们提出nnMobileNet++,一种混合架构,逐步融合卷积与变压器表示。该框架包含三个核心组件:(i) 动态蛇形卷积用于边界感知特征提取;(ii) 在第二次下采样后引入阶段特定的变压器块以建模全局上下文;(iii) 眼底图像预训练以提升泛化能力。在多个公开眼底数据集上的分类实验及消融研究表明,nnMobileNet++在保持低计算成本的同时,达到或接近当前最优性能,凸显其作为轻量高效眼底图像分析框架的巨大潜力。

原文摘要 · Abstract (English)

Retinal imaging is a critical, non-invasive modality for the early detection and monitoring of ocular and systemic diseases. Deep learning, particularly convolutional neural networks (CNNs), has significant progress in automated retinal analysis, supporting tasks such as fundus image classification, lesion detection, and vessel segmentation. As a representative lightweight network, nnMobileNet has demonstrated strong performance across multiple retinal benchmarks while remaining computationally efficient. However, purely convolutional architectures inherently struggle to capture long-range dependencies and model the irregular lesions and elongated vascular patterns that characterize on retinal images, despite the critical importance of vascular features for reliable clinical diagnosis. To further advance this line of work and extend the original vision of nnMobileNet, we propose nnMobileNet++, a hybrid architecture that progressively bridges convolutional and transformer representations. The framework integrates three key components: (i) dynamic snake convolution for boundary-aware feature extraction, (ii) stage-specific transformer blocks introduced after the second down-sampling stage for global context modeling, and (iii) retinal image pretraining to improve generalization. Experiments on multiple public retinal datasets for classification, together with ablation studies, demonstrate that nnMobileNet++ achieves state-of-the-art or highly competitive accuracy while maintaining low computational cost, underscoring its potential as a lightweight yet effective framework for retinal image analysis.

眼底图像轻量模型混合架构医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。