arXiv:2412.00943cs.LG2024-12被引 1

从可解释性视角重新审视模型校准,发现简单决策树效果不逊于复杂方法。

Calibration through the Lens of Interpretability

  • 提出校准的公理化框架,系统梳理理想校准模型的性质
  • 实验证明简单可解释决策树在多数指标上媲美主流校准方法
  • 适合关注模型可信度与可解释性的研究人员参考

校准是获取可靠标签概率估计的关键,要求模型输出值准确反映真实标签概率。然而校准并不保证分类精度,也不一定具备可解释性,且在有限数据下难以验证。现有众多评估指标和损失函数,各自衡量校准的不同方面。本文首次对校准概念进行公理化研究,系统归纳理想校准模型的期望属性及其对应评估指标,并分析其可行性与关联性。同时通过实证评估,比较常见校准方法与一种简单可解释的决策树方法,结果表明后者在多数场景下表现相当甚至更优。

原文摘要 · Abstract (English)

Calibration is a frequently invoked concept when useful label probability estimates are required on top of classification accuracy. A calibrated model is a function whose values correctly reflect underlying label probabilities. Calibration in itself however does not imply classification accuracy, nor human interpretable estimates, nor is it straightforward to verify calibration from finite data. There is a plethora of evaluation metrics (and loss functions) that each assess a specific aspect of a calibration model. In this work, we initiate an axiomatic study of the notion of calibration. We catalogue desirable properties of calibrated models as well as corresponding evaluation metrics and analyze their feasibility and correspondences. We complement this analysis with an empirical evaluation, comparing common calibration methods to employing a simple, interpretable decision tree.

模型校准可解释性决策树评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。