arXiv:2608.24691cs.AIcs.CL2026-08

大模型在信息不全时高估自己判断,导致行动失误

Confident at the moment of action: belief miscalibration in LLM play under hidden information

论文配图:Confident at the moment of action: belief miscalibration in LLM play under hidden information
图 1 · 摘自论文原文
  • 通过隐藏王位的国际象棋变体测试模型信念校准度
  • 高置信度判断正确率仅1.6%(62次中1次)
  • 传统评估指标无法发现信念偏差,适合需可信决策的研究者

智能体越来越多地根据自身置信度决定行动,这假设置信度与正确性在行动时刻一致。我们在一个隐藏信息的国际象棋变体中检验这一假设:王的位置可秘密、反复转移至不同棋子。每回合单独提取模型对对手隐藏王位置的概率分布,并在赛后与真实情况对比。在两个独立批次中,置信度≥0.5时做出的捕获行动,仅1/62正确。校准缺陷几乎全部集中于此类事件:原始批次占99.3%,复现批次占98.7%。该模式在四种跨厂商模型配置中仍一致出现(点估计),虽统计差异不显著,但同一模型在固定外部排行榜分数下,仅改变推理预算即可使指标变化幅度接近大模型间差距。此外,传统评估维度(合法性、成本、延迟、完成率)与信念质量完全脱节,表现最佳的配置反而信念质量最差。模型即使信念严重失准,仍可能赢得比赛,因此仅看结果的评估无法察觉此问题。

原文摘要 · Abstract (English)

Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverable after the game. Across two independent batches, captures made at high stated confidence ($\geq 0.5$) about the hidden piece's location were correct in 1 of 62 cases. The calibration deficit is concentrated almost entirely in these events: 99.3% of it in the original batch, 98.7% in the replication. The same pattern, in weaker form, orders consistently (point estimates only; most pairwise gaps are not statistically distinguishable at this sample size) across four further model configurations spanning a second provider -- reported as scope for the finding, not as evidence that capability predicts calibration: a same-model comparison at a fixed external leaderboard score shows a deliberation-budget change alone moves the metric by nearly as much as a large cross-model gap. In a separate seat, conventional evaluation axes -- legality, cost, latency, completion rate -- can dissociate entirely from belief quality, with the configuration winning on every conventional axis producing the worst belief quality tested. A model exhibiting this pattern can still win the game its belief was about, which is why outcome-only evaluation would not detect it.

大模型信念校准博弈推理评估漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。