arXiv:2502.06243cs.CVcs.LG2025-02被引 20

多尺度Transformer提升皮肤病变分类准确率

Multi-Scale Transformer Architecture for Accurate Medical Image Classification

  • 设计多尺度特征融合与注意力优化机制
  • 在ISIC 2017上超越ResNet50等主流模型
  • 可视化显示模型关注区域与病灶匹配

本研究提出一种基于改进Transformer架构的AI皮肤病变分类算法,解决医学图像分析中精度与鲁棒性挑战。通过引入多尺度特征融合机制并优化自注意力过程,模型能有效提取全局与局部特征,增强对边界模糊、结构复杂的病灶检测能力。在ISIC 2017数据集上的评估显示,该模型在准确率、AUC、F1-Score和精确率等关键指标上均优于ResNet50、VGG19、ResNext及Vision Transformer。Grad-CAM可视化进一步验证了模型可解释性,其关注区域与实际病灶高度一致。研究展示了先进AI模型在医学影像中的变革潜力,为更精准可靠的诊断工具提供支持。未来将探索该方法在更广泛医学影像任务中的可扩展性,并研究多模态数据融合以强化智能医疗诊断框架。

原文摘要 · Abstract (English)

This study introduces an AI-driven skin lesion classification algorithm built on an enhanced Transformer architecture, addressing the challenges of accuracy and robustness in medical image analysis. By integrating a multi-scale feature fusion mechanism and refining the self-attention process, the model effectively extracts both global and local features, enhancing its ability to detect lesions with ambiguous boundaries and intricate structures. Performance evaluation on the ISIC 2017 dataset demonstrates that the improved Transformer surpasses established AI models, including ResNet50, VGG19, ResNext, and Vision Transformer, across key metrics such as accuracy, AUC, F1-Score, and Precision. Grad-CAM visualizations further highlight the interpretability of the model, showcasing strong alignment between the algorithm's focus areas and actual lesion sites. This research underscores the transformative potential of advanced AI models in medical imaging, paving the way for more accurate and reliable diagnostic tools. Future work will explore the scalability of this approach to broader medical imaging tasks and investigate the integration of multimodal data to enhance AI-driven diagnostic frameworks for intelligent healthcare.

医学图像Transformer皮肤病变多尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。