arXiv:2508.09381cs.CVcs.AI2025-08被引 5

研究皮肤病变分割中人工标注差异,发现恶性程度越高差异越大。

What Can We Learn from Inter-Annotator Variability in Skin Lesion Segmentation?

  • 分析多标注者数据,揭示恶性病变标注一致性更低
  • 仅用皮肤镜图像就能预测标注差异,误差仅0.108
  • 将标注一致性作软特征提升模型性能,准确率提高4.2%

医学图像分割因边界模糊、标注者偏好、经验与工具差异等因素存在标注者内和跨标注者变异。边界模糊的病变(如毛刺状或浸润性结节,或符合ABCD法则的不规则边缘)尤其易引发分歧,且常与恶性相关。本文构建了目前最大的多标注者皮肤病变分割数据集IMA++,深入研究了标注者、恶性程度、工具和技能等因素导致的变异。发现跨标注者一致性(IAA,以Dice衡量)与病变恶性程度存在显著关联(p<0.001)。进一步证明,仅从皮肤镜图像即可准确预测IAA,平均绝对误差为0.108。最后,将IAA作为“软”临床特征引入多任务学习目标,使多种模型架构在IMA++及四个公开皮肤镜数据集上平衡准确率平均提升4.2%。代码已开源。

原文摘要 · Abstract (English)

Medical image segmentation exhibits intra- and inter-annotator variability due to ambiguous object boundaries, annotator preferences, expertise, and tools, among other factors. Lesions with ambiguous boundaries, e.g., spiculated or infiltrative nodules, or irregular borders per the ABCD rule, are particularly prone to disagreement and are often associated with malignancy. In this work, we curate IMA++, the largest multi-annotator skin lesion segmentation dataset, on which we conduct an in-depth study of variability due to annotator, malignancy, tool, and skill factors. We find a statistically significant (p<0.001) association between inter-annotator agreement (IAA), measured using Dice, and the malignancy of skin lesions. We further show that IAA can be accurately predicted directly from dermoscopic images, achieving a mean absolute error of 0.108. Finally, we leverage this association by utilizing IAA as a "soft" clinical feature within a multi-task learning objective, yielding a 4.2% improvement in balanced accuracy averaged across multiple model architectures and across IMA++ and four public dermoscopic datasets. The code is available at https://github.com/sfu-mial/skin-IAV.

皮肤病变标注差异多任务学习医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。