伪造的证据让大模型盲目自信,即使全假也照样下判断。
Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
- 用虚构数据包装问题,显著提升模型承诺回答率
- 虚假面板使承诺率从24.5%升至36.8%,与真实数据无差别
- 模型错在行动决策机制,而非理解或判断能力
当大模型看到一个专业外观的市场面板时,其对不可预测问题做出方向性判断的意愿,相较直接提问提升了近48个百分点——12个前沿模型中,承诺率从6.5%飙升至54.0%。即使面板上所有数据均为虚构,模型仍保持36.8%的承诺率,与真实数据(37.6%)统计无异。真正驱动信心的是信息的呈现形式而非内容真实性。模型在可答问题上准确率接近完美,且对问题不可知性的判断正确率达90%,但在此类问题上仅0.4%选择行动。失败点在于“是否行动”的决策门控机制,且该机制可在特定条件下通过监督微调修复。针对540个合成案例(如骰子、硬币、罐子、计时器)进行微调后,30亿参数模型在原任务上承诺率降至0.0%,并成功泛化至三个新领域。但若强制使用固定格式,模型会失去推理空间,反而在本可正确回答的问题上自信犯错。决策门控可训练但依赖上下文,部署需同时兼顾两者。
原文摘要 · Abstract (English)
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。