arXiv:2512.16030cs.AI2025-12被引 4

测试大模型对未知事件的自信程度是否靠谱,发现全都不准。

Do Large Language Models Know What They Don't Know? Kalshibench: A New Benchmark for Evaluating Epistemic Calibration via Prediction Markets

  • 用真实未来事件预测题测模型对不确定性的判断能力
  • 所有模型都严重高估自信,最好也差得很远(ECE=0.12)
  • 更强推理反而更不靠谱,适合评估模型可信度的研究者看

一个校准良好的模型应使其信心水平与实际准确率一致——当它声称80%把握时,应真正正确80%。尽管大语言模型在各类任务上表现卓越,但其认知校准能力仍不清楚。我们引入卡尔希基准(KalshiBench),包含300个来自监管级预测市场平台卡尔希(Kalshi)的真实未来事件问题,其结果在模型训练截止后才揭晓。不同于传统静态知识测试,该基准评估模型对真正未知未来的不确定性量化能力。我们评测了五款前沿模型:Claude Opus 4.5、GPT-5.2、DeepSeek-V3.2、Qwen3-235B 和 Kimi-K2,发现所有模型均存在系统性过度自信。即使表现最佳的模型(Claude Opus 4.5,ECE=0.120)仍存在显著校准误差,而增强推理能力的模型如GPT-5.2-XHigh校准更差(ECE=0.395),尽管准确率相当。关键的是,仅有一款模型获得正向布里尔技能得分,表明多数模型的表现不如直接使用基础概率。结果表明,模型规模和推理增强并不自动带来校准优势,凸显认知校准是一种需专门开发的能力。

原文摘要 · Abstract (English)

A well-calibrated model should express confidence that matches its actual accuracy -- when it claims 80\% confidence, it should be correct 80\% of the time. While large language models (LLMs) have achieved remarkable performance across diverse tasks, their epistemic calibration remains poorly understood. We introduce \textbf{KalshiBench}, a benchmark of 300 prediction market questions from Kalshi, a CFTC-regulated exchange, with verifiable real-world outcomes occurring after model training cutoffs. Unlike traditional benchmarks measuring accuracy on static knowledge, KalshiBench evaluates whether models can appropriately quantify uncertainty about genuinely unknown future events. We evaluate five frontier models -- Claude Opus 4.5, GPT-5.2, DeepSeek-V3.2, Qwen3-235B, and Kimi-K2 -- and find \textbf{systematic overconfidence across all models}. Even the best-calibrated model (Claude Opus 4.5, ECE=0.120) shows substantial calibration errors, while reasoning-enhanced models like GPT-5.2-XHigh exhibit \emph{worse} calibration (ECE=0.395) despite comparable accuracy. Critically, only one model achieves a positive Brier Skill Score, indicating most models perform worse than simply predicting base rates. Our findings suggest that scaling and enhanced reasoning do not automatically confer calibration benefits, highlighting epistemic calibration as a distinct capability requiring targeted development.

大模型校准不确定性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。