测试大模型在物理题中的不确定性和准确率关系,发现推理越难不确定性越明显。
Testing Uncertainty of Large Language Models for Physics Knowledge and Reasoning
- 通过多选物理题测试模型回答的准确性与不确定性关联
- 模型在高置信度时准确率高,但整体呈钟形分布且不普适
- 复杂逻辑推理任务中准确率与不确定性差异更显著
近年来,大型语言模型(LLMs)在多个领域问答中表现出色,但其存在“幻觉”现象,难以评估预测可靠性。本文分析了主流开源LLM及gpt-3.5 Turbo在多项选择物理题上的表现,聚焦答案准确率与物理相关主题不确定性的关系。研究发现,多数模型在高置信度时能给出准确答案,但这一行为并非普遍规律。准确率与不确定性之间的关系呈现广泛的水平钟形分布。随着题目对逻辑推理要求提高,准确率与不确定性的不对称性加剧;而在知识检索类任务中,该关系依然明显。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have gained significant popularity in recent years for their ability to answer questions in various fields. However, these models have a tendency to "hallucinate" their responses, making it challenging to evaluate their performance. A major challenge is determining how to assess the certainty of a model's predictions and how it correlates with accuracy. In this work, we introduce an analysis for evaluating the performance of popular open-source LLMs, as well as gpt-3.5 Turbo, on multiple choice physics questionnaires. We focus on the relationship between answer accuracy and variability in topics related to physics. Our findings suggest that most models provide accurate replies in cases where they are certain, but this is by far not a general behavior. The relationship between accuracy and uncertainty exposes a broad horizontal bell-shaped distribution. We report how the asymmetry between accuracy and uncertainty intensifies as the questions demand more logical reasoning of the LLM agent, while the same relationship remains sharp for knowledge retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。