arXiv:2506.21883cs.CV2025-06被引 4

用梯度分析自动识别皮肤影像中的虚假模式,提升远程诊断准确率

GRASP-PsONet: Gradient-based Removal of Spurious Patterns for PsOriasis Severity Classification

  • 通过梯度追踪定位导致模型误判的异常训练图像
  • 剔除8.2%问题图像后,模型AUC-ROC提升5个百分点至90%
  • 可高效发现标注不一致样本,减少人工审评成本

银屑病(PsO)严重程度评分对临床试验至关重要,但受评估者间差异和现场评估负担限制。远程成像利用患者手机拍摄照片具有可扩展性,但存在光照、背景、设备质量等难以察觉的变异,影响模型性能。这些因素与皮肤病专家标注不一致共同降低了自动化评分的可靠性。本文提出一种基于梯度的可解释性方法,自动识别引入虚假相关性的训练图像,从而改善模型泛化能力。通过追踪误分类验证图像的梯度,检测与标注不一致或受细微非临床伪影影响的训练样本。该方法应用于基于ConvNeXT的弱监督模型,用于从手机图像中分类银屑病严重程度。剔除8.2%标记图像后,模型在保留测试集上的AUC-ROC从85%提升至90%。通常需多位专家评审以保证标注准确性,成本高耗时。本方法能有效识别标注不一致样本,在仅审查前30%样本情况下,检出超过90%的评估者间分歧。显著提升远程评估的自动化评分鲁棒性,应对数据采集中的多样性挑战。

原文摘要 · Abstract (English)

Psoriasis (PsO) severity scoring is important for clinical trials but is hindered by inter-rater variability and the burden of in person clinical evaluation. Remote imaging using patient captured mobile photos offers scalability but introduces challenges, such as variation in lighting, background, and device quality that are often imperceptible to humans but can impact model performance. These factors, along with inconsistencies in dermatologist annotations, reduce the reliability of automated severity scoring. We propose a framework to automatically flag problematic training images that introduce spurious correlations which degrade model generalization, using a gradient based interpretability approach. By tracing the gradients of misclassified validation images, we detect training samples where model errors align with inconsistently rated examples or are affected by subtle, nonclinical artifacts. We apply this method to a ConvNeXT based weakly supervised model designed to classify PsO severity from phone images. Removing 8.2% of flagged images improves model AUC-ROC by 5% (85% to 90%) on a held out test set. Commonly, multiple annotators and an adjudication process ensure annotation accuracy, which is expensive and time consuming. Our method detects training images with annotation inconsistencies, potentially removing the need for manual review. When applied to a subset of training data rated by two dermatologists, the method identifies over 90% of cases with inter-rater disagreement by reviewing only the top 30% of samples. This improves automated scoring for remote assessments, ensuring robustness despite data collection variability.

皮肤病弱监督可解释性远程医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。