提出评估概率分类器可信度的新框架,用局部校准衡量模型是否值得信赖。
I-trustworthy Models. A framework for trustworthiness evaluation of probabilistic classifiers
- 基于局部校准误差设计假设检验方法,量化模型可信度。
- 引入核方法计算局部校准误差(KLCE),并给出收敛性理论保证。
- 可诊断模型偏差,适用于医疗、金融等高风险场景的可信评估。
随着概率模型在社会各个领域日益普及并推动科学进步,仅依赖预测准确率等传统指标已不足,必须评估其可信度。本文基于信任的能力理论,提出I-trustworthy框架——一种通过将局部校准与可信度关联,评估概率分类器在推断任务中可信度的新方法。为评估I-trustworthiness,采用局部校准误差(LCE)并开发了一种基于核的假设检验方法,使用核局部校准误差(KLCE)检验模型的局部校准性能。研究提供了KLCE无偏估计量的收敛性边界,并设计诊断工具以识别和度量校准偏差。该方法在模拟数据和真实数据集上均验证了有效性。同时分析了多种再校准方法的LCE表现,证明现有方法无法充分实现I-trustworthiness。
原文摘要 · Abstract (English)
As probabilistic models continue to permeate various facets of our society and contribute to scientific advancements, it becomes a necessity to go beyond traditional metrics such as predictive accuracy and error rates and assess their trustworthiness. Grounded in the competence-based theory of trust, this work formalizes I-trustworthy framework -- a novel framework for assessing the trustworthiness of probabilistic classifiers for inference tasks by linking local calibration to trustworthiness. To assess I-trustworthiness, we use the local calibration error (LCE) and develop a method of hypothesis-testing. This method utilizes a kernel-based test statistic, Kernel Local Calibration Error (KLCE), to test local calibration of a probabilistic classifier. This study provides theoretical guarantees by offering convergence bounds for an unbiased estimator of KLCE. Additionally, we present a diagnostic tool designed to identify and measure biases in cases of miscalibration. The effectiveness of the proposed test statistic is demonstrated through its application to both simulated and real-world datasets. Finally, LCE of related recalibration methods is studied, and we provide evidence of insufficiency of existing methods to achieve I-trustworthiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。