arXiv:2607.03882cs.CLcs.AI2026-07

大模型能一致描述风险,但常错估概率大小。

Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language

论文配图:Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
图 1 · 摘自论文原文
  • 用贝塔分布模拟预测,让大模型为概率值配文字描述
  • 模型对可能性描述准确,对不确定性评估偏差严重
  • 即使给预计算数据,仍无法解决文字表达失准问题

大模型越来越多地被用作AI输出的后验解释器,但其在自然语言中传达概率信息的可靠性仍不明确。为胜任此角色,模型需对相同输入给出一致描述,并选择与数值大小匹配的表述。我们评估了九个大模型在两阶段预测流程中的表现:上游模型生成带有似然和不确定性的概率输出,大模型则为其选择合适的文字描述。通过采样贝塔分布(以众数和先验样本量参数化)模拟预测,并在六个领域场景和十种温度设置下进行实验,每项重复十次。结果发现,大模型整体具有一致性但存在显著失准,尤其在不确定性任务上表现更差。提供预计算的统计量(众数和先验样本量)虽降低了上下文干扰,但未能解决根本性失准问题,表明瓶颈在于文字生成环节本身。当前大模型尚不能作为零样本独立的风险沟通工具用于概率预测。

原文摘要 · Abstract (English)

LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language. For this role to be viable, models must produce identical verbal descriptions for identical inputs, and select descriptions that accurately reflect the magnitude of the underlying numerical quantities. We evaluate whether nine LLMs meet these requirements within a two-stage prediction pipeline, in which an upstream model has produced probabilistic outputs characterized by their likelihood and uncertainty, and LLMs are tasked with selecting an appropriate verbal descriptor for each. We simulate predictions from an upstream model by taking samples from a Beta distribution parameterized by its mode and prior sample size. We then prompt LLMs to explain these predictions under six domain contexts and with ten temperature settings, and repeating each experiment ten times. We find that LLMs are generally consistent but miscalibrated, with substantially weaker performance on uncertainty than on likelihood tasks. Providing models with precomputed summary statistics (mode and prior sample size) reduced sensitivity to contextual framing but did not resolve the underlying miscalibration, suggesting that the bottleneck resides in the verbalization step itself. These findings indicate that current LLMs do not yet constitute reliable zero-shot standalone risk communication tools for probabilistic predictions.

大模型评估风险沟通概率推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。