提出ACUTE协议,让大模型更可信地表达不确定性和决策价值。
The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust

- 基于模型激活值设计通用的置信度与可信度评估方法
- 在6个模型上实现比基线更高的信息量与校准性平衡
- 适合需高可信决策的场景,如医疗、金融问答
随着大语言模型能力提升并广泛部署,可信度变得至关重要。校准性是信任的良好代理:良好的置信度估计有助于权衡信任特定输出的风险与收益。然而,即便模型不断进步,其置信度仍普遍偏高,存在过度自信问题。此外,校准可被操纵:一种始终预测基础率的策略虽完全校准,但毫无信息量。为此,我们提出新指标——由最优参考者归一化的期望效用(EURO),在兼顾校准性与信息量之间取得平衡。同时,我们开发了通用的基于激活值的置信度、效用与可信度估计协议(ACUTE),适用于多项选择题回答、工具调用和科学文档摘要等3类任务,在6个来自4个模型家族的模型上表现优异。ACUTE在EURO指标上超越强基线,且保持低校准误差。结果表明,为大模型配备ACUTE协议可显著提升其在多种场景下的校准性、效用与可信度。
原文摘要 · Abstract (English)
As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is a good proxy for trust: well-calibrated confidence estimates help inform the risk versus reward tradeoff when trusting a specific model output. Unfortunately, even as models improve, they remain poorly calibrated, often biasing towards overconfidence. Additionally, calibration can be gamed: a policy that always predicts the base rate is perfectly calibrated, but completely uninformative. To resolve this, we develop a new metric, expected utility renormalized by the oracle (EURO), that balances calibration and informativeness. We also propose a general-purpose activation-based confidence, utility, and trust estimation protocol (ACUTE) to appropriately adjudicate uncertainty. The ACUTE protocol provides flexible, sample-efficient, and compute-efficient confidence estimators for 3 tasks including multiple choice question answering, tool-calling, and scientific document summarization across 6 models from 4 model families. ACUTE outperforms strong baselines on EURO, while maintaining low calibration error. Taken together, our work shows that equipping LLMs with the ACUTE protocol can improve calibration, utility, and trustworthiness in numerous settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。