评测大模型在真实世界不确定性下的推理能力,发现其预测常不准确且过于自信。
OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World Data
- 设计多领域数值估计任务,让模型输出概率先验来评估不确定性推理。
- 六款前沿模型的先验分布普遍不准确且过度自信,性能提升有限。
- 适合关注模型可信度、风险决策与真实场景应用的研究者参考。
实际应用中,语言模型需处理信息不全和不确定性问题,但现有评估多聚焦于答案明确的任务。由于难以设计出模型难答而人类能答对的难题,模型在不确定性推理方面表现仍不清晰。为此,我们提出 OpenEstimate,一个可扩展的多领域基准,用于评估模型在需整合大量背景知识并输出概率先验的数值估计任务中的表现。我们评估这些先验的准确性和校准性,并与真实分布样本对比。在六款前沿语言模型上测试发现,模型生成的先验通常不准确且过于自信。尽管不同提示策略影响有限,但响应方式对性能有小幅改善。OpenEstimate 为先进模型提供了挑战性评估,也为开发更好不确定性推理模型提供平台。
原文摘要 · Abstract (English)
Real-world settings where language models (LMs) are deployed -- in domains spanning healthcare, finance, and other forms of knowledge work -- require models to grapple with incomplete information and reason under uncertainty. Yet most LM evaluations focus on problems with well-defined answers and success criteria. This gap exists in part because natural problems involving uncertainty are difficult to construct: given that LMs have access to most of the same knowledge as humans, it is non-trivial to design questions for which LMs will struggle to produce correct answers, but which humans can answer reliably. As a result, LM performance on reasoning under uncertainty remains poorly characterized. To address this gap, we introduce OpenEstimate, an extensible, multi-domain benchmark for evaluating LMs on numerical estimation tasks that require models to synthesize significant amounts of background information and express predictions as probabilistic priors. We assess these priors for accuracy and calibration, quantifying their usefulness relative to samples from the true distribution of interest. Across six frontier LMs, we find that LM-elicited priors are often inaccurate and overconfident. Performance improves modestly depending on how uncertainty is elicited from the model, but is largely unaffected by changes in sampling strategy, reasoning effort, or prompt design. The OpenEstimate benchmark thus offers a challenging evaluation for frontier LMs and a platform for developing models that are better at probabilistic estimation and reasoning under uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。