LLM的自信表达与真实决策脱节,高风险下也不愿放弃回答。
Are LLM Decisions Faithful to Verbal Confidence?
- 设计新框架RiskEval,测试模型在不同错误代价下的应对策略。
- 即使高惩罚下应频繁放弃回答,模型仍几乎从不拒绝,导致效用崩溃。
- 当前模型缺乏将不确定感转化为合理决策的能力,可信系统需更深层策略
大型语言模型(LLMs)能生成看似复杂的自身不确定性估计。然而,这种表达出的信心是否真正关联到模型的推理、知识或决策仍不明确。为此,我们提出$ extbf{RiskEval}$:一个评估模型是否根据误差惩罚变化调整弃权策略的框架。对多个前沿模型的评估显示关键脱节:模型在表达口头信心时并不具备成本意识,且在高惩罚条件下无法战略性地选择是否参与或放弃。即使极端惩罚使频繁弃权成为数学最优策略,模型几乎从不弃权,导致效用严重下降。这表明,仅靠校准的口头信心评分不足以构建可信、可解释的AI系统,因为当前模型缺乏将不确定性信号转化为最优、风险敏感决策的战略能力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can produce surprisingly sophisticated estimates of their own uncertainty. However, it remains unclear to what extent this expressed confidence is tied to the reasoning, knowledge, or decision making of the model. To test this, we introduce $\textbf{RiskEval}$: a framework designed to evaluate whether models adjust their abstention policies in response to varying error penalties. Our evaluation of several frontier models reveals a critical dissociation: models are neither cost-aware when articulating their verbal confidence, nor strategically responsive when deciding whether to engage or abstain under high-penalty conditions. Even when extreme penalties render frequent abstention the mathematically optimal strategy, models almost never abstain, resulting in utility collapse. This indicates that calibrated verbal confidence scores may not be sufficient to create trustworthy and interpretable AI systems, as current models lack the strategic agency to convert uncertainty signals into optimal and risk-sensitive decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。