arXiv:2410.02210cs.CLcs.LG2024-10

让大模型学会区分正确与错误预测的可信度。

Calibrate to Discriminate: Improve In-Context Learning with Label-Free Comparative Inference

  • 设计新方法通过对比推理提升模型判断力
  • 在5个数据集上显著改善预测准确率与置信度一致性
  • 适合需要可靠置信度评估的零样本/少样本场景

尽管大语言模型(LLMs)的上下文学习表现优异,我们发现其存在一种独特的校准偏差:无论预测正确与否,模型均给出相同置信度。这种现象称为无差别校准偏差。传统校准指标如期望校准误差(ECE)难以有效捕捉该行为。为此,我们提出新指标以衡量无差别校准偏差的严重程度,并开发了一种新型上下文对比推理方法,缓解校准偏差并提升分类性能。在五个数据集上的大量实验表明,相比常规零样本与少样本提示,该方法可实现更准确且校准更优的预测结果。

原文摘要 · Abstract (English)

While in-context learning with large language models (LLMs) has shown impressive performance, we have discovered a unique miscalibration behavior where both correct and incorrect predictions are assigned the same level of confidence. We refer to this phenomenon as indiscriminate miscalibration. We found that traditional calibration metrics, such as Expected Calibrated Errors (ECEs), are unable to capture this behavior effectively. To address this issue, we propose new metrics to measure the severity of indiscriminate miscalibration. Additionally, we develop a novel in-context comparative inference method to alleviate miscalibrations and improve classification performance. Through extensive experiments on five datasets, we demonstrate that our proposed method can achieve more accurate and calibrated predictions compared to regular zero-shot and few-shot prompting.

大模型校准上下文学习推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。