arXiv:2608.14768cs.CVcs.LG2026-08

用不确定性识别皮肤病变难分类样本,提升临床可信度。

Uncertainty Identifies Difficult Samples Across Methods: A Multi-Task Study on a Heterogeneous Skin Lesion Dataset

  • 统一骨干网络+双任务头,融合多来源皮肤数据集
  • 五种不确定量化方法中,深度集成表现最优且样本难易排名一致
  • 高不确定性样本集中错误,适合用于选择性转诊

皮肤病变分类器在关键病例上可能自信地出错,因此判断预测是否可信与预测本身同样重要。我们基于来自多个ISIC来源的数据集,采用共享骨干网络和两个联合学习的分支:二分类(恶性vs良性)与五分类诊断分支。比较了五种不确定性量化方法(MC Dropout、DropConnect、Flipout、Deep Ensembles、DUQ)在准确率、校准性、不确定性分解和风险-覆盖率上的表现。结果表明,困难样本的判定具有高度方法无关性:即使熵分布较窄的方法,对同一组样本的难易排序也高度一致(单样本熵相关系数0.54至0.91)。方法选择更影响校准性和不确定性分解,其中深度集成表现最佳;而发现难样本的能力则普遍良好,使最不确定样本的延迟决策可移除大量错误,支持基于不确定性的选择性转诊,本研究仅评估了分布内情况。

原文摘要 · Abstract (English)

Skin lesion classifiers can be confidently wrong on the cases that matter most, so knowing when a prediction should not be trusted is clinically as useful as the prediction. We study uncertainty quantification on a dataset pooled from many ISIC sources, with a shared backbone and two jointly learned heads: a binary malignant versus non-malignant head and a five-class diagnostic head. Five UQ methods (MC Dropout, DropConnect, Flipout, Deep Ensembles, DUQ) are compared on accuracy, calibration, uncertainty decomposition, and risk-coverage. Difficulty is largely method-agnostic: even methods with narrow entropy distributions rank the same samples as hard (per-sample entropy correlations of $0.54$ to $0.91$). The choice of method matters more for calibration and uncertainty decomposition, where Deep Ensembles is the clear winner, than for finding difficult cases. The ranking is also good enough that deferring the most uncertain cases removes a disproportionate share of errors, supporting uncertainty-based selective referral, evaluated here in-distribution only.

皮肤病变不确定性深度集成选择性转诊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。