arXiv:2409.12180cs.CLcs.LG2024-09被引 26

让大模型学会用恰当语气表达不确定,提升可信度。

Finetuning Language Models to Emit Linguistic Expressions of Uncertainty

  • 用模型自身置信度指导微调,生成更准确的不确定表述。
  • 在多个问答数据集上验证,单条答案的不确定性表达显著改善。
  • 适合需要可信赖输出的医疗、金融等高风险场景使用。

大型语言模型(LLMs)在信息检索和决策任务中应用日益广泛。尽管用途广泛,但它们生成的内容常与现实事实冲突,且其说服性表达容易使错误信息显得自信而可信。这导致用户难以将模型表达的自信程度与其预测准确性对齐,进而产生盲目信任或全盘否定两种极端。本文探索了基于不确定性增强预测的监督微调方法,以训练模型生成校准后的不确定性语言表达。具体而言,我们首先评估预训练模型的校准能力,再通过其自身置信度进行微调,使其生成更可靠的不确定性表述。在多个问答数据集上的实验表明,LLMs在预测评估上具有良好的校准性,基于自身信心的监督微调能有效生成校准良好的不确定性表达,尤其在单条主张回答中表现突出。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly employed in information-seeking and decision-making tasks. Despite their broad utility, LLMs tend to generate information that conflicts with real-world facts, and their persuasive style can make these inaccuracies appear confident and convincing. As a result, end-users struggle to consistently align the confidence expressed by LLMs with the accuracy of their predictions, often leading to either blind trust in all outputs or a complete disregard for their reliability. In this work, we explore supervised finetuning on uncertainty-augmented predictions as a method to develop models that produce linguistic expressions of uncertainty. Specifically, we measure the calibration of pre-trained models and then fine-tune language models to generate calibrated linguistic expressions of uncertainty. Through experiments on various question-answering datasets, we demonstrate that LLMs are well-calibrated in assessing their predictions, and supervised finetuning based on the model's own confidence leads to well-calibrated expressions of uncertainty, particularly for single-claim answers.

大模型不确定性微调可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。