测试大模型对事实、虚构与预测的认知能力,发现其不确定性表达不可靠。
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
- 用可控故事测试大模型的语义认知能力
- 模型生成的不确定表达不一致且不可靠
- 需增强模型对认识论模态的语义理解
理性说话者应清楚自己知道什么、不知道什么,并根据证据强度生成相应表述。然而,当前大语言模型在不确定现实环境中仍难以生成与信心水平匹配的表达。尽管近期通过言语化不确定性来估计和校准模型置信度已成趋势,但缺乏对模型隐空间中编码的不确定性语言知识的深入考察。本文基于认识论表达的类型学框架,利用受控故事评估大模型对认识论情态的知识。实验表明,大模型在生成认识论表达方面表现有限且不稳健,因此其生成的不确定性表达并非始终可靠。要构建具备不确定性意识的大模型,必须丰富其对认识论情态的语义知识。
原文摘要 · Abstract (English)
Rational speakers are supposed to know what they know and what they do not know, and to generate expressions matching the strength of evidence. In contrast, it is still a challenge for current large language models to generate corresponding utterances based on the assessment of facts and confidence in an uncertain real-world environment. While it has recently become popular to estimate and calibrate confidence of LLMs with verbalized uncertainty, what is lacking is a careful examination of the linguistic knowledge of uncertainty encoded in the latent space of LLMs. In this paper, we draw on typological frameworks of epistemic expressions to evaluate LLMs' knowledge of epistemic modality, using controlled stories. Our experiments show that the performance of LLMs in generating epistemic expressions is limited and not robust, and hence the expressions of uncertainty generated by LLMs are not always reliable. To build uncertainty-aware LLMs, it is necessary to enrich semantic knowledge of epistemic modality in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。