arXiv:2506.10119cs.CVcs.AI2025-06

对比CNN与视觉变换器在银屑病图像分类中的表现,发现小模型的ViT更高效。

Detecção da Psoríase Utilizando Visão Computacional: Uma Abordagem Comparativa Entre CNNs e Vision Transformers

  • 用预训练模型适配银屑病图像数据集,比较CNN与视觉变换器
  • 双注意力视觉变换器-基线(DaViT-B)达96.4%的F1分数,最优
  • 适合医疗图像分类研究者参考,尤其关注轻量化模型设计

本文比较了卷积神经网络(CNNs)与视觉变换器(ViTs)在多类分类银屑病及类似疾病皮损图像任务中的表现。采用ImageNet预训练模型进行微调。两类模型均取得高预测性能,但视觉变换器在小模型规模下表现更优。双注意力视觉变换器-基线(DaViT-B)获得最佳结果,F1分数达96.4%,推荐作为自动化银屑病检测中最高效的架构。研究强化了视觉变换器在医学图像分类任务中的潜力。

原文摘要 · Abstract (English)

This paper presents a comparison of the performance of Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) in the task of multi-classifying images containing lesions of psoriasis and diseases similar to it. Models pre-trained on ImageNet were adapted to a specific data set. Both achieved high predictive metrics, but the ViTs stood out for their superior performance with smaller models. Dual Attention Vision Transformer-Base (DaViT-B) obtained the best results, with an f1-score of 96.4%, and is recommended as the most efficient architecture for automated psoriasis detection. This article reinforces the potential of ViTs for medical image classification tasks.

银屑病检测视觉变换器医学图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。