发现大模型在事实类问题上拥有自有的正确性判断能力
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness

- 用模型自身隐藏状态与外部模型对比判断答案正确性
- 在意见分歧的数据集上,自表示在事实题中显著更准
- 这种优势出现在早期到中期层,暗示特定记忆检索机制
人类通过无法被外部观察的内在状态评估理解程度。我们探究大语言模型是否具备类似的、关于答案正确性的私有知识。通过在模型自身隐藏状态和外部模型表示上训练正确性分类器,测试自表示是否具有性能优势。在标准评测中未发现优势:自探针表现与同行模型探针相当。我们推测这是由于模型间答案正确性高度一致所致。为分离真正的私有知识,我们在模型产生矛盾预测的分歧子集上进行评估。结果显示:在事实类任务中,自表示持续优于同行表示;但在数学推理中无明显优势。进一步分析显示,事实类优势从早期到中期层逐步显现,符合模型特异性记忆检索特征;而数学推理在任意层级均无稳定优势。
原文摘要 · Abstract (English)
Humans use introspection to evaluate their understanding through private internal states inaccessible to external observers. We investigate whether large language models possess similar privileged knowledge about answer correctness, information unavailable through external observation. We train correctness classifiers on question representations from both a model's own hidden states and external models, testing whether self-representations provide a performance advantage. On standard evaluation, we find no advantage: self-probes perform comparably to peer-model probes. We hypothesize this is due to high inter-model agreement of answer correctness. To isolate genuine privileged knowledge, we evaluate on disagreement subsets, where models produce conflicting predictions. Here, we discover domain-specific privileged knowledge: self-representations consistently outperform peer representations in factual knowledge tasks, but show no advantage in math reasoning. We further localize this domain asymmetry across model layers, finding that the factual advantage emerges progressively from early-to-mid layers onward, consistent with model-specific memory retrieval, while math reasoning shows no consistent advantage at any depth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。