arXiv:2511.15476cs.CVcs.AI2025-11被引 4

融合CNN与Transformer,提升猴痘皮肤图像检测精度

RS-CA-HSICT: A Residual and Spatial Channel Augmented CNN Transformer Framework for Monkeypox Detection

  • 设计新型混合架构,结合残差、空间与通道增强机制
  • 在两个数据集上达98.30%准确率,超越传统CNN与ViT
  • 适合医学图像分析与皮肤病智能诊断研究者

本文提出一种混合深度学习框架RS-CA-HSICT,结合卷积神经网络(CNN)与视觉变压器(Transformer)优势,用于提升猴痘(MPox)检测性能。该框架包含HSICT模块、残差CNN模块、空间CNN块和通道增强(CA)模块,有效增强多尺度特征空间、病变细节及长程依赖关系。新设计的HSICT模块整合了茎干CNN与定制的ICT块,实现高效多头注意力与统一(H)及结构(S)操作;其中H/S层分别学习空间同质性与精细结构细节,抑制噪声并建模复杂形态变化。逆残差学习缓解梯度消失,阶段式分辨率降低保障尺度不变性。此外,通过TL驱动的残差与空间CNN图谱对学习到的HSICT通道进行增强,捕捉全局与局部结构线索、细微纹理及对比度变化。通道融合与注意力模块筛选判别性通道,抑制冗余信息,提升计算效率。最后,空间注意力机制优化像素选择,精准识别猴痘中的微弱模式与类内对比差异。在Kaggle基准数据集和一个多样化的猴痘数据集上的实验表明,分类准确率达到98.30%,F1-score为98.13%,显著优于现有CNN与ViT模型。

原文摘要 · Abstract (English)

This work proposes a hybrid deep learning approach, namely Residual and Spatial Learning based Channel Augmented Integrated CNN-Transformer architecture, that leverages the strengths of CNN and Transformer towards enhanced MPox detection. The proposed RS-CA-HSICT framework is composed of an HSICT block, a residual CNN module, a spatial CNN block, and a CA, which enhances the diverse feature space, detailed lesion information, and long-range dependencies. The new HSICT module first integrates an abstract representation of the stem CNN and customized ICT blocks for efficient multihead attention and structured CNN layers with homogeneous (H) and structural (S) operations. The customized ICT blocks learn global contextual interactions and local texture extraction. Additionally, H and S layers learn spatial homogeneity and fine structural details by reducing noise and modeling complex morphological variations. Moreover, inverse residual learning enhances vanishing gradient, and stage-wise resolution reduction ensures scale invariance. Furthermore, the RS-CA-HSICT framework augments the learned HSICT channels with the TL-driven Residual and Spatial CNN maps for enhanced multiscale feature space capturing global and localized structural cues, subtle texture, and contrast variations. These channels, preceding augmentation, are refined through the Channel-Fusion-and-Attention block, which preserves discriminative channels while suppressing redundant ones, thereby enabling efficient computation. Finally, the spatial attention mechanism refines pixel selection to detect subtle patterns and intra-class contrast variations in Mpox. Experimental results on both the Kaggle benchmark and a diverse MPox dataset reported classification accuracy as high as 98.30% and an F1-score of 98.13%, which outperforms the existing CNNs and ViTs.

猴痘检测混合模型医学图像注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。