arXiv:2511.13240cs.LG2025-11

大模型自信与行为不一致,知之未必行之。

Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions

  • 通过交互场景测试模型信心与行为的一致性
  • 高自信时反而违背自身预测,低自信时不调用工具
  • 静态校准不能保证动态决策可靠,适合评估者关注

大型语言模型(LLMs)在代理和多轮交互任务中应用日益广泛,其行动后果重大。为可靠部署并管理风险,需获取模型的不确定性估计。然而,当前信心提取方法多在静态数据集(如问答基准)上评估,未直接检验其在交互场景中的表现。本文研究静态信心估计与动态行为之间的关系,发现显著的“行为-信念差距”:模型常做出与其信心相悖的行动。在预测市场中,模型常反向押注自身高置信预测;在工具使用中,低信心时未能调用信息查询工具;在用户挑战中,高自信时改变答案,而低自信时坚持原答。关键发现是,静态校准强弱无法预测动态一致性——更强、更校准的模型有时比小型开源模型更不一致。结果揭示当前评估方法的重大盲点:知道自己的知识不等于会据此理性行动。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in agentic and multi-turn workflows where they are tasked to perform actions of significant consequence. In order to deploy them reliably and manage risky outcomes in these settings, it is helpful to access model uncertainty estimates. However, confidence elicitation methods for LLMs are typically not evaluated directly in agentic settings; instead, they are evaluated on static datasets, such as Q&A benchmarks. In this work we investigate the relationship between confidence estimates elicited in static settings and the behavior of LLMs in interactive settings. We uncover a significant action-belief gap -- LLMs frequently take actions that contradict their elicited confidences. In a prediction market setting, we find that models often bet against their own high-confidence predictions; in a tool-use setting, models fail to reliably invoke information-seeking tools when their internal confidence is low; and in a user-challenge setting, models change their answers when they have high confidence in them, whilst sticking to answers they have low confidence in. Crucially, we show that static calibration is an insufficient predictor of consistency in the above dynamic settings, as stronger, better calibrated models are somtimes less consistent than their smaller and weaker open-source counterparts. Our results highlight a critical blind spot in current evaluation methodologies: ensuring that a model knows what it knows does not guarantee that it will act rationally on that knowledge.

大模型信心评估行为一致性交互系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。