arXiv:2509.23000cs.LGcs.DS2025-09被引 1

提出新校准方法,用少量样本实现多分类模型精准输出概率分布。

Sample-efficient Multiclass Calibration under $\ell_{p}$ Error

  • 设计插值型校准误差定义,兼顾样本效率与理论保证。
  • 仅需多项式样本即可校准任意多分类模型,接近最优误差依赖性。
  • 创新使用高自适应数据分析,显著降低样本开销,适合实际部署。

多分类预测器输出标签分布时,校准尤为困难,因其预测值数量呈指数增长。本文提出一种新的校准误差定义,介于两种已有校准误差度量之间:一种具有已知的指数样本复杂度,另一种具有多项式样本复杂度。所提算法可对任意给定预测器在插值范围内的绝大多数情况完成校准,仅在端点之一例外,且仅需多项式样本量。在另一端点,其误差参数依赖性近乎最优,优于先前工作。关键技术贡献在于首次将高自适应数据分析应用于该问题,虽自适应性强,但样本复杂度仅增加对数级开销。

原文摘要 · Abstract (English)

Calibrating a multiclass predictor, that outputs a distribution over labels, is particularly challenging due to the exponential number of possible prediction values. In this work, we propose a new definition of calibration error that interpolates between two established calibration error notions, one with known exponential sample complexity and one with polynomial sample complexity for calibrating a given predictor. Our algorithm can calibrate any given predictor for the entire range of interpolation, except for one endpoint, using only a polynomial number of samples. At the other endpoint, we achieve nearly optimal dependence on the error parameter, improving upon previous work. A key technical contribution is a novel application of adaptive data analysis with high adaptivity but only logarithmic overhead in the sample complexity.

多分类校准样本效率统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。