仅用一次查询即可精准审计大模型的校准度
Single-Query Black-Box Calibration Auditing via Logit Bias

- 利用logit_bias参数实现单次查询精确测量概率阈值
- 提出可证明一致的二分类校准误差估计器
- 适合需要高效评估黑箱大模型的开发者
评估大语言模型(LLMs)的校准度对于其作为零样本分类器的安全部署至关重要。然而,越来越多的商业API提供商隐藏了标准校准指标所需的连续输出概率。为突破这一透明性限制,我们证明,任何暴露logit_bias参数的LLM API均可通过数学操控,仅用一次查询即可确定精确的概率阈值。基于此机制,我们提出一种全新的、可证明一致的二分类任务真校准误差估计器。该方法因此提供了一种高效的黑箱基础模型审计框架。
原文摘要 · Abstract (English)
Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the continuous output probabilities required by standard calibration metrics. To bypass this opacity, we demonstrate that any LLM API exposing a logit\_bias parameter can be mathematically manipulated to evaluate exact probability thresholds using strictly one query per sample. Leveraging this mechanism, we introduce a novel and provably consistent estimator of the True Calibration Error for binary tasks. Our approach therefore provides an efficient framework for auditing black-box foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。