用虚拟赌局让大模型暴露预测信心,让模糊判断变清晰。
Going All-In on LLM Accuracy: Fake Prediction Markets, Real Confidence Signals
- 把模型评估改成虚拟赌局,用下注金额反映预测信心。
- 大额下注准确率高达99%,小额下注仅74%。
- 虽准确率提升不显著,但首次让模型自信程度可见可用。
大型语言模型在评估其他模型时通常缺乏置信度表示。本试点研究测试了将评估任务设计为赌局(虚构的预测市场,使用自有LLM币)是否能提升预测准确性并揭示校准后的置信信号。我们生成了100道可验证答案的数学与逻辑题。六种基线模型(三类当前代、三类前代)回答所有题目。三种预测模型则对每道题-基线组合预测其正确性。每个预测模型在两种条件下完成匹配实验:对照组(简单正确/错误预测)和激励组(预测加下注1-100,000 LLMCoin,起始资金1,000,000 LLMCoin,赔率均等)。每种条件共5,400次预测。激励组准确率略高(81.5% vs. 79.1%,p = .089,d = 0.86),且学习速度显著更快(第1到第4轮提升12.0个百分点,对照组仅2.9个百分点,p = .011)。最显著的是,下注金额与信心高度相关:40,000+大额下注正确率达约99%,小额下注(<1,000)准确率仅约74%。关键发现并非虚拟货币使模型更聪明——准确率提升未达统计显著性(p = .089);而是赌局机制创造了二元输出中缺失的可读置信信号。这表明简单的金融框架可能使大模型成为风险感知型预测者,使其内部信念可见可被利用。该协议为未来元评估系统及大模型间预测市场奠定基础。
原文摘要 · Abstract (English)
Large language models are increasingly used to evaluate other models, yet these judgments typically lack any representation of confidence. This pilot study tests whether framing an evaluation task as a betting game (a fictional prediction market with its own LLM currency) improves forecasting accuracy and surfaces calibrated confidence signals. We generated 100 math and logic questions with verifiable answers. Six Baseline models (three current-generation, three prior-generation) answered all items. Three Predictor models then forecasted, for each question-baseline pair, if the baseline would answer correctly. Each predictor completed matched runs in two conditions: Control (simple correct/incorrect predictions) and Incentive (predictions plus wagers of 1-100,000 LLMCoin under even odds, starting from a 1,000,000 LLMCoin bankroll). Across 5,400 predictions per condition, Incentive runs showed modestly higher accuracy (81.5% vs. 79.1%, p = .089, d = 0.86) and significantly faster learning across rounds (12.0 vs. 2.9 percentage-point improvement from Round 1 to Round 4, p = .011). Most notably, stake size tracked confidence. "Whale" bets of 40,000+ coins were correct ~99% of the time, while small bets (<1,000 coins) showed only ~74% accuracy. The key finding is not that fictional money makes models smarter; accuracy gains were modest and did not reach statistical significance (p = .089) in this pilot. Rather, the betting mechanic created a legible confidence signal absent from binary yes/no outputs. This suggests that simple financial framing may help transform LLMs into risk-aware forecasters, making their internal beliefs visible and usable. The protocol offers a foundation for future work for meta-evaluation systems and what may become LLM-to-LLM prediction markets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。