探究大模型不确定性与人类的相似性,发现其行为与内部激活模式中存在类人不确定性信号。
Human-Alignment, Calibration, and Activation Patterns in Large Language Model Uncertainty

- 通过对比模型输出与内部激活模式,检测类人不确定性信号
- 多数据集验证下模型同时具备不确定性对齐与校准能力
- 指令微调会显著影响不确定性对齐表现,适合关注可信AI的研究者
不确定性量化是大语言模型行为分析的重要分支,主要为识别和抑制幻觉,聚焦于校准能力——即模型对不确定性的判断是否准确对应任务表现。本文首次系统探讨大模型不确定性与人类的相似性,研究其外部行为及内部激活模式中是否存在类人不确定性信号(称为不确定性对齐)。在涵盖多项选择与开放式事实回忆的多个数据集上,验证了模型是否同时具备不确定性对齐与校准特性,并分析了指令微调对这两方面的影响。
原文摘要 · Abstract (English)
Uncertainty Quantification is a large and growing subfield of large language model behavioral analysis. Primarily to recognize and combat hallucination, the field has largely focused on measuring and improving calibration, the accuracy of uncertainty judgments to task efficacy. In this work, we investigate the relatively underexplored question of how similar large language model uncertainty is to human uncertainty. We investigate the presence and strength of human-similar uncertainty signals, deemed uncertainty alignment, in large language model overt behavior and internal activation patterns. We identify whether the models show evidence of simultaneous alignment and calibration on a variety of datasets covering both multiple choice and open ended factual recall. And we characterize the effect of instruct fine-tuning on each of these facets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。