模型越自信越可能胡说,但通过分析生成过程可有效识别假话。
Epistemic Observability in Language Models

- 用熵和概率分布等计算痕迹替代仅看文本,突破不可观测困局。
- 每令牌熵的检测准确率达0.757,比纯文本方法高2.5至3.9个百分点。
- 为系统设计者提供验证资源分配的实用决策地图,适合可信AI建设者。
我们发现,模型在捏造信息时报告的信心最高。在四个模型家族(OLMo-3、Llama-3.1、Qwen3、Mistral)中,自报告置信度与准确性呈负相关,AUC值介于0.28至0.36之间,低于随机猜测的0.5。我们在严格形式假设下证明,这不是能力缺陷,而是观测限制所致。仅通过文本观察(即监督者只能看到模型输出文本)时,任何监控系统都无法可靠区分真实输出与合理虚构。我们证明:第一,仅依赖输入查询的策略无法在模糊世界状态下实现认知诚实;第二,任何优化来自文本监督者奖励的学习算法,当真实与虚构响应在监督者眼中观测相同,则无法收敛至诚实行为。在我们的形式化模型中,这些不可能性不随模型规模或训练方式(包括RLHF与指令微调)改变。我们构建了张量接口,导出与正确性结构耦合的计算副产品(如每令牌熵和对数概率分布),突破上述不可能性。每令牌熵的综合AUC达0.757,在所有测试预算水平(10%、20%、30%)上均优于所有文本基线2.5–3.9个百分点。该熵信号在不同架构间具有泛化能力(Spearman ρ = 0.762)。核心贡献是成本曲面——将验证预算(接受昂贵检查的查询比例)与每种裁判策略的检测准确率之间的经验映射,为系统构建者提供实际资源分配参考。贡献在于这张地图,真正的领土是你正在构建的系统。
原文摘要 · Abstract (English)
We find that models report highest confidence precisely when they are fabricating. Across four model families (OLMo-3, Llama-3.1, Qwen3, Mistral), self-reported confidence inversely correlates with accuracy, with AUC ranging from 0.28 to 0.36 where 0.5 is random guessing. We prove, under explicit formal assumptions, that this is not a capability gap but an observational one. Under text-only observation, where a supervisor sees only the model's output text, no monitoring system can reliably distinguish honest model outputs from plausible fabrications. We prove two results: first, that any policy conditioning only on the query cannot satisfy epistemic honesty across ambiguous world states; second, that no learning algorithm optimizing reward from a text-only supervisor can converge to honest behavior when the supervisor's observations are identical for both grounded and fabricated responses. Within our formal model, these impossibilities hold regardless of model scale or training procedure, including RLHF and instruction tuning. We construct a tensor interface that escapes the impossibility by exporting computational byproducts (per-token entropy and log-probability distributions) that are structurally coupled to correctness under standard training. Per-token entropy achieves pooled AUC 0.757, outperforming all text baselines by 2.5--3.9 percentage points at every budget level tested (10\%, 20\%, 30\%). The entropy signal generalizes across architectures (Spearman $ρ= 0.762$). The core contribution is a cost surface where the empirical mapping from verification budget (fraction of queries receiving expensive checks) to detection accuracy for each judge strategy is a practical lookup for system builders deciding how to allocate verification resources. The contribution is the map. The territory is the system you are building.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。