细粒度校准让大模型长文本生成更可信
Atomic Calibration of LLMs in Long-Form Generations
- 将长文本拆成原子事实,逐条评估置信度
- 发现大模型在长文本中置信度校准差
- 揭示置信度变化规律,指导未来改进
大语言模型在长文本生成中常出现幻觉,严重制约其实际应用。置信度校准作为幻觉的有效指标,对提升模型可信度至关重要。以往研究多关注短文本的单次输出置信度(宏观校准),难以应对包含真伪混杂信息的长文本。本文系统研究原子级校准,将长文本分解为原子事实,进行细粒度的事实性置信度评估。我们进一步将现有置信度获取方法分为判别式与生成式两类,并提出两种新的置信度融合策略以提升校准效果。实验表明,大语言模型在长文本生成中的原子级校准表现较差。更重要的是,原子校准揭示了不同置信度方法间的对齐模式及置信度随生成过程的变化规律,为未来长文本置信度估计研究提供了新方向。
原文摘要 · Abstract (English)
Large language models (LLMs) often suffer from hallucinations, posing significant challenges for real-world applications. Confidence calibration, as an effective indicator of hallucination, is thus essential to enhance the trustworthiness of LLMs. Prior work mainly focuses on short-form tasks using a single response-level score (macro calibration), which is insufficient for long-form outputs that may contain both accurate and inaccurate claims. In this work, we systematically study atomic calibration, which evaluates factuality calibration at a fine-grained level by decomposing long responses into atomic claims. We further categorize existing confidence elicitation methods into discriminative and generative types, and propose two new confidence fusion strategies to improve calibration. Our experiments demonstrate that LLMs exhibit poorer calibration at the atomic level during long-form generation. More importantly, atomic calibration uncovers insightful patterns regarding the alignment of confidence methods and the changes of confidence throughout generation. This sheds light on future research directions for confidence estimation in long-form generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。