让大模型生成更真实:根据输入动态调整事实性校准
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
- 基于自适应校准,根据提示动态调整事实性判断
- 在多个任务中显著提升条件覆盖度,减少错误生成
- 支持选择性预测,可自动过滤不可靠内容,适合高可靠性场景
大语言模型容易生成事实错误的内容。已有方法使用近似预测提供事实性的不确定性估计和统计保证,但通常不随提示变化,难以捕捉输入依赖的差异性,导致在某些任务或提示下过滤过少(过度覆盖)或过多(覆盖不足)。本文提出一种自适应近似预测方法,扩展了近似评分变换技术以适配大语言模型,应用于长文本生成和多选题问答。该方法实现提示依赖的校准,在保持边际覆盖保证的同时提升条件覆盖。此外,该方法天然支持选择性预测,可在下游应用中过滤不可靠陈述或答案选项。我们在多个白盒模型和多样领域上评估该方法,结果表明其在条件覆盖方面显著优于现有基线。
原文摘要 · Abstract (English)
Large language models (LLMs) are prone to generating factually incorrect outputs. Recent work has applied conformal prediction to provide uncertainty estimates and statistical guarantees for the factuality of LLM generations. However, existing approaches are typically not prompt-adaptive, limiting their ability to capture input-dependent variability. As a result, they may filter out too few items (leading to over-coverage) or too many (under-coverage) for a given task or prompt. We propose an adaptive conformal prediction approach that extends conformal score transformation methods to LLMs, with applications to long-form generation and multiple-choice question answering. This enables prompt-dependent calibration, retaining marginal coverage guarantees while improving conditional coverage. In addition, the approach naturally supports selective prediction, allowing unreliable claims or answer choices to be filtered out in downstream applications. We evaluate our approach on multiple white-box models across diverse domains and show that it significantly outperforms existing baselines in terms of conditional coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。