让大模型学会准确表达不确定性的方法,提升判断可靠性。
Generalization of Fine-Tuned Uncertainty Communication and Metacognition in Large Language Models
- 通过多任务微调训练模型,提高自信度与正确率的一致性
- 在数学、常识等任务上,正确答案的置信度显著更高
- 跨任务迁移有限,联合训练能更好泛化
大型语言模型在需要可靠不确定性表达的场景中越来越重要,但其自述置信度常与实际正确性不一致。本文测试了监督微调是否能改善不确定性表达,并评估效果在不同领域和任务形式间的迁移能力。对两个模型在通用知识、数学和开放问答题上进行微调,评估单题置信度估计(报告单一答案的数值置信度)和成对置信度比较(判断哪个问题更可能答对)。在训练领域外的医学、法律和真实性基准测试中评估校准度、区分度和准确率。结果表明,微调显著提升了置信度与正确性的对齐,使模型更倾向于为正确答案赋予更高置信度;该改进在训练领域内明显,在新领域中较弱;但单一任务训练无法可靠地在单题置信度估计与成对比较间迁移。多任务微调在所研究模型和任务中带来更广泛的提升。结论:大模型的不确定性表达可通过训练提升,但跨元认知任务的泛化能力有限。联合训练多个置信度任务或可支持更广泛的推广,但仍需在更多模型和任务上验证。
原文摘要 · Abstract (English)
Background. Large language models are increasingly used in settings where confident but incorrect answers can mislead users. Reliable uncertainty communication requires a form of metacognition: monitoring when one's own answers are likely to be correct. Yet models' stated confidence is often poorly aligned with answer correctness. We test whether supervised fine-tuning improves uncertainty communication and whether gains transfer across domains and task formats. Methods. We fine-tuned two models on general knowledge, mathematics, and open-ended trivia questions. We evaluated single-question confidence estimation, in which the model reports numeric confidence for one answer, and pairwise confidence comparison, in which it chooses which of two questions it is more likely to answer correctly. We tested held-out questions from training domains and new medical, legal, and truthfulness benchmarks. We assessed calibration, discrimination, and answer accuracy before and after fine-tuning. Results. Here we show that fine-tuning improves alignment between stated confidence and observed accuracy and increases the model's ability to assign higher confidence to correct than to incorrect answers. Gains occur within training domains and, to a lesser extent, in new domains. However, single-task training does not reliably transfer between single-question confidence estimation and pairwise confidence comparison. Multitask fine-tuning produces broader gains in the models and tasks studied here. Conclusions. Uncertainty communication in large language models is trainable, but transfer across metacognitive tasks is limited. Joint training on multiple confidence tasks may support broader generalization, although further tests across model families and metacognitive tasks are needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。