arXiv:2510.05566stat.MLcs.AI2025-10中稿 · ICML被引 6

让大模型在领域变化下仍能可靠判断自己是否犯错

Domain-Shift-Aware Conformal Prediction for Large Language Models

  • 根据测试样本与校准数据的相似度重加权,适应领域变化
  • 在MMLU上验证,领域偏移时覆盖率显著优于传统方法
  • 适合对可靠性要求高的真实场景部署

大语言模型在诸多任务中表现优异,但常产生过度自信且事实错误的输出(即幻觉),在实际应用中带来风险。置信预测可提供有限样本、无需分布假设的覆盖保证,但在领域偏移下会失效,导致覆盖率不足且预测集不可靠。本文提出领域偏移感知置信预测(DS-CP)框架,通过系统性地根据校准样本与测试提示的接近程度进行重加权,使置信预测在大模型领域偏移下仍保持有效性并提升适应性。理论分析与在MMLU基准上的实验表明,该方法在显著分布偏移下比标准置信预测更可靠,同时保持高效,为大模型在真实部署中的可信不确定性量化提供了实用路径。

原文摘要 · Abstract (English)

Large language models have achieved impressive performance across diverse tasks. However, their tendency to produce overconfident and factually incorrect outputs, known as hallucinations, poses risks in real-world applications. Conformal prediction provides finite-sample, distribution-free coverage guarantees, but standard conformal prediction breaks down under domain shift, often leading to under-coverage and unreliable prediction sets. We propose a new framework called Domain-Shift-Aware Conformal Prediction (DS-CP). Our framework adapts conformal prediction to large language models under domain shift, by systematically reweighting calibration samples based on their proximity to the test prompt, thereby preserving validity while enhancing adaptivity. Our theoretical analysis and experiments on the MMLU benchmark demonstrate that the proposed method delivers more reliable coverage than standard conformal prediction, especially under substantial distribution shifts, while maintaining efficiency. This provides a practical step toward trustworthy uncertainty quantification for large language models in real-world deployment.

置信预测大模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。