arXiv:2502.19684cs.CL2025-02ACL被引 8

构建人类与模型对比的校准评估基准,精准测量大模型推理信心可靠性。

GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration

  • 设计渐进线索问答任务,按揭示顺序评估模型回答时机与信心
  • 1749组人类与模型竞赛数据表明:人类虽准确率低但更可信
  • 提出CalScore指标,可识别模型偏离人类校准模式的错误类型

语言模型常出现过度自信的错误回答。本文提出GRACE——一个结合人类校准对比的语言模型校准评估基准。该基准包含一系列问题,每个问题由逐步变易的线索构成,最终导向同一答案;模型需在线索逐渐披露时尽早正确作答。此设置可精细衡量模型在何时、多准、多稳地给出答案。我们通过实时举办人机对抗比赛,收集了1,749组人类与模型团队的答题时间、准确率和信心数据。提出CalScore指标,利用GRACE分析模型校准误差,并识别出与人类行为不同的模型校准偏差类型。结果发现,尽管模型准确率高于人类,但人类整体校准更优。当前最先进模型在GRACE上表现不佳,证明其有效评估模型校准改进进展。

原文摘要 · Abstract (English)

Language models are often miscalibrated, leading to confidently incorrect answers. We introduce GRACE, a benchmark for language model calibration that incorporates comparison with human calibration. GRACE consists of question-answer pairs, in which each question contains a series of clues that gradually become easier, all leading to the same answer; models must answer correctly as early as possible as the clues are revealed. This setting permits granular measurement of model calibration based on how early, accurately, and confidently a model answers. After collecting these questions, we host live human vs. model competitions to gather 1,749 data points on human and model teams' timing, accuracy, and confidence. We propose a metric, CalScore, that uses GRACE to analyze model calibration errors and identify types of model miscalibration that differ from human behavior. We find that although humans are less accurate than models, humans are generally better calibrated. Since state-of-the-art models struggle on GRACE, it effectively evaluates progress on improving model calibration.

模型校准评估基准人类对比推理可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。