让大模型回答的置信度更可信,不同问题类型下都准。
QA-Calibration of Language Model Confidence Scores
- 按问题类别分组校准置信度,比整体平均更可靠。
- 在多个LLM和数据集上验证,校准后准确率提升显著。
- 无需假设分布,适合医疗、金融等高风险场景使用。
为使生成式问答系统用于决策或关键应用,其返回的置信度必须真实反映答案正确性。现有校准方法仅保证置信度在平均意义上与答案正确性一致,但难以支持具体决策。本文提出针对生成式QA的新型校准标准——QA-calibration,要求校准在不同问题-答案组中均成立。为此设计了离散化的后处理校准方案,并提供无需分布假设的性能保障。在多个QA基准和大型语言模型(LLMs)上,通过提示词提取的置信度验证了该方法的有效性。
原文摘要 · Abstract (English)
To use generative question-and-answering (QA) systems for decision-making and in any critical application, these systems need to provide well-calibrated confidence scores that reflect the correctness of their answers. Existing calibration methods aim to ensure that the confidence score is, *on average*, indicative of the likelihood that the answer is correct. We argue, however, that this standard (average-case) notion of calibration is difficult to interpret for decision-making in generative QA. To address this, we generalize the standard notion of average calibration and introduce QA-calibration, which ensures calibration holds across different question-and-answer groups. We then propose discretized posthoc calibration schemes for achieving QA-calibration. We establish distribution-free guarantees on the performance of this method and validate our method on confidence scores returned by elicitation prompts across multiple QA benchmarks and large language models (LLMs).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。