arXiv:2409.12390cs.CV2024-09中稿 · WACV2025被引 13

用多模态融合与关联建模提升皮肤病变多标签分类准确率

A Novel Perspective for Multi-modal Multi-label Skin Lesion Classification

  • 设计三模态跨注意力架构,融合图像与患者信息实现深层特征融合
  • 在Derm7pt数据集上达到77.85%诊断准确率,优于现有最优方法
  • 特别适合需要处理多标签、不平衡数据的医学图像分析任务

基于深度学习的皮肤疾病辅助诊断(CAD)方法的性能依赖于对多种数据模态(临床+皮肤镜图像及患者元数据)的分析,并解决多标签分类难题。现有方法通常局限于有限的多模态技术,将多标签问题视为多个多分类问题,忽视了学习不平衡和标签相关性问题。本文提出创新的皮肤病变分类器,采用基于Transformer的多模态多标签模型(SkinM2Former)。针对多模态分析,引入三模态跨注意力Transformer(TMCT),在Transformer编码器的不同特征层级融合三类模态(图像与元数据)。针对多标签分类,设计多头注意力模块以学习标签间相关性,并优化解决多标签与不平衡学习问题。在公开的Derm7pt数据集上,SkinM2Former实现77.27%的平均分类准确率和77.85%的平均诊断准确率,超越现有最先进方法。

原文摘要 · Abstract (English)

The efficacy of deep learning-based Computer-Aided Diagnosis (CAD) methods for skin diseases relies on analyzing multiple data modalities (i.e., clinical+dermoscopic images, and patient metadata) and addressing the challenges of multi-label classification. Current approaches tend to rely on limited multi-modal techniques and treat the multi-label problem as a multiple multi-class problem, overlooking issues related to imbalanced learning and multi-label correlation. This paper introduces the innovative Skin Lesion Classifier, utilizing a Multi-modal Multi-label TransFormer-based model (SkinM2Former). For multi-modal analysis, we introduce the Tri-Modal Cross-attention Transformer (TMCT) that fuses the three image and metadata modalities at various feature levels of a transformer encoder. For multi-label classification, we introduce a multi-head attention (MHA) module to learn multi-label correlations, complemented by an optimisation that handles multi-label and imbalanced learning problems. SkinM2Former achieves a mean average accuracy of 77.27% and a mean diagnostic accuracy of 77.85% on the public Derm7pt dataset, outperforming state-of-the-art (SOTA) methods.

皮肤病变多模态多标签Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。