arXiv:2507.06867stat.MLcs.CV2025-07被引 7

解决长尾分类中预测集过大或覆盖不足的难题

Conformal Prediction for Long-Tailed Classification

  • 提出预估调整的softmax得分函数,优化各类别平均覆盖率
  • 通过线性插值融合边际与类别条件覆盖,平衡集合大小与覆盖率
  • 在Pl@ntNet-300K和iNaturalist-2018上验证有效性

许多现实中的分类任务(如植物识别)存在极端长尾类分布。为使预测集在该场景下实用,需同时满足:(i) 保证各类别条件覆盖,避免稀有类别被系统性遗漏;(ii) 预测集大小合理,便于人工验证。然而,现有符合性预测方法在长尾设置下迫使实践者在小集合但覆盖差,或大集合但覆盖好之间二选一。本文提出具有边际覆盖保证的方法,可平滑权衡集合大小与类别条件覆盖。首先引入新的符合性得分函数——预估调整的softmax,以优化宏观覆盖(即各类别条件覆盖的平均值)。其次提出一种新流程,通过线性插值边际与类别条件符合性预测的得分阈值实现折衷。我们在包含1,081类和8,142类的Pl@ntNet-300K与iNaturalist-2018两个长尾图像数据集上验证了所提方法。

原文摘要 · Abstract (English)

Many real-world classification problems, such as plant identification, have extremely long-tailed class distributions. In order for prediction sets to be useful in such settings, they should (i) provide good class-conditional coverage, ensuring that rare classes are not systematically omitted from the prediction sets, and (ii) be a reasonable size, allowing users to easily verify candidate labels. Unfortunately, existing conformal prediction methods, when applied to the long-tailed setting, force practitioners to make a binary choice between small sets with poor class-conditional coverage or sets that have very good class-conditional coverage but are extremely large. We propose methods with marginal coverage guarantees that smoothly trade off set size and class-conditional coverage. First, we introduce a new conformal score function called prevalence-adjusted softmax that optimizes for macro-coverage, defined as the average class-conditional coverage across classes. Second, we propose a new procedure that interpolates between marginal and class-conditional conformal prediction by linearly interpolating their conformal score thresholds. We demonstrate our methods on Pl@ntNet-300K and iNaturalist-2018, two long-tailed image datasets with 1,081 and 8,142 classes, respectively.

长尾分类符合性预测图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。