arXiv:2503.01688cs.CLcs.LG2025-03被引 5

研究大模型回答的不确定性,发现熵值可判断知识类问题难易度。

When an LLM is apprehensive about its answers -- and when its uncertainty is justified

  • 用词元熵和模型自评评估不同题型的不确定性
  • 生物类问题熵值预测错误率AUC达0.73,数学类仅为0.55
  • 现有数据集需平衡推理量,避免评估偏差

不确定性估计对评估大语言模型至关重要,尤其在高风险场景中。本文研究了在14个不同主题的多选题任务中,词元熵与模型作为裁判(MASJ)两种方法的表现。实验涵盖三个不同规模的模型(Phi-4、Mistral、Qwen,1.5B至72B)。结果表明,MASJ表现接近随机猜测;而响应熵在知识依赖型领域(如生物)能有效预测错误,其ROC AUC为0.73;但在推理依赖型领域(如数学),相关性消失(数学类ROC-AUC为0.55)。进一步发现,熵值估计需依赖模型推理量,因此应在不确定性框架中整合数据不确定性;同时,现有MMLU-Pro数据集存在推理量偏差,需平衡各子领域的推理需求以实现公平评估。

原文摘要 · Abstract (English)

Uncertainty estimation is crucial for evaluating Large Language Models (LLMs), particularly in high-stakes domains where incorrect answers result in significant consequences. Numerous approaches consider this problem, while focusing on a specific type of uncertainty, ignoring others. We investigate what estimates, specifically token-wise entropy and model-as-judge (MASJ), would work for multiple-choice question-answering tasks for different question topics. Our experiments consider three LLMs: Phi-4, Mistral, and Qwen of different sizes from 1.5B to 72B and $14$ topics. While MASJ performs similarly to a random error predictor, the response entropy predicts model error in knowledge-dependent domains and serves as an effective indicator of question difficulty: for biology ROC AUC is $0.73$. This correlation vanishes for the reasoning-dependent domain: for math questions ROC-AUC is $0.55$. More principally, we found out that the entropy measure required a reasoning amount. Thus, data-uncertainty related entropy should be integrated within uncertainty estimates frameworks, while MASJ requires refinement. Moreover, existing MMLU-Pro samples are biased, and should balance required amount of reasoning for different subdomains to provide a more fair assessment of LLMs performance.

不确定性估计大模型评估知识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。