arXiv:2505.23709cs.CVcs.AI2025-05被引 5

通过多模态对比学习,提升皮肤病变图像的表型识别能力。

Skin Lesion Phenotyping via Nested Multi-modal Contrastive Learning

  • 构建嵌套对比学习框架,融合病变图像与患者级元数据。
  • 在分类任务中显著优于现有预训练方法,提升识别准确率。
  • 适合皮肤病学研究与临床辅助诊断系统开发人员使用。

我们提出SLIMP(Skin Lesion Image-Metadata Pre-training),通过一种新颖的嵌套对比学习方法,捕捉图像与元数据之间的复杂关系,以学习皮肤病变的丰富表征。仅依赖图像进行黑色素瘤检测和皮肤病变分类面临巨大挑战,主要源于成像条件(光照、色彩、分辨率、距离等)的巨大差异,以及缺乏临床和表型上下文信息。临床医生通常结合患者的病史及其他病变外观,采取整体评估方式判断病变风险及是否需要切除。受此启发,SLIMP将个体病变的外观特征与患者层面的医疗记录及其它临床相关信息相结合。通过在整个学习过程中充分利用所有可用的数据模态,该预训练策略在下游皮肤病变分类任务中表现优于其他预训练方法,凸显了所学表征的质量。

原文摘要 · Abstract (English)

We introduce SLIMP (Skin Lesion Image-Metadata Pre-training) for learning rich representations of skin lesions through a novel nested contrastive learning approach that captures complex relationships between images and metadata. Melanoma detection and skin lesion classification based solely on images, pose significant challenges due to large variations in imaging conditions (lighting, color, resolution, distance, etc.) and lack of clinical and phenotypical context. Clinicians typically follow a holistic approach for assessing the risk level of the patient and for deciding which lesions may be malignant and need to be excised, by considering the patient's medical history as well as the appearance of other lesions of the patient. Inspired by this, SLIMP combines the appearance and the metadata of individual skin lesions with patient-level metadata relating to their medical record and other clinically relevant information. By fully exploiting all available data modalities throughout the learning process, the proposed pre-training strategy improves performance compared to other pre-training strategies on downstream skin lesions classification tasks highlighting the learned representations quality.

皮肤病变多模态学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。