发现语言模型自我判断会干扰正确性探测,导致结果误判。
Diagnosing Correctness Probes under Self-Judgement Confounding
- 构造矛盾案例,分离模型自评与客观正确性信号
- 实验证明自评信号比客观正确信号更具可迁移性
- 适用于检测模型自我认知偏差的研究者和安全评估人员
隐藏状态读出可预测语言模型输出的正确性,但客观正确性(OC)通常与模型自身自评(SJ)一致,使解码信号语义模糊。本文构建了OC与SJ预测相反的冲突案例。在高置信度分歧情况下,传统以正确性标注的对比常将错误但自我认可的回答排在正确但被拒绝的回答之上,遵循自评而非客观正确性。我们估计了与自评和客观正确性相关的方向,并在数学推理与事实回忆任务中评估其极性。在四个指令微调模型(最大140亿参数)中,自评相关方向在跨领域迁移中均显著优于随机水平,而客观正确性相关方向在所有对应条件下点估计值低于随机水平。该不对称性在中后期层发展,且在答案似然、序列长度及零方向控制下仍成立,扩展至MMLU和二分类TruthfulQA任务无需目标领域方向拟合。在所研究模型与诊断子集上,最可靠可迁移的成分始终保留自评方向极性。因此,迁移性本身无法确立客观正确性语义。
原文摘要 · Abstract (English)
Hidden-state readouts can predict whether language-model outputs are correct, but objective correctness (OC) usually agrees with the model's own self-judgement (SJ), leaving the decoded signal semantically ambiguous. We construct conflict cases in which OC and SJ predict opposite readout orderings. On high-confidence disagreements, conventional correctness-labelled contrasts often rank incorrect/self-endorsed responses above correct/self-rejected responses, following SJ rather than OC. We estimate factorial SJ- and OC-associated directions and evaluate their polarity across mathematical reasoning and factual recall. Across four instruction-tuned models up to 14B parameters, the SJ-associated direction transfers above chance in both cross-domain directions for every model, whereas the OC-associated direction has a below-chance point estimate for the expected OC ordering in every corresponding condition. This transfer asymmetry develops across middle-to-late layers, persists under answer-likelihood, sequence-length, and null-direction controls, and extends to MMLU and binary TruthfulQA without target-domain direction fitting. Across the studied models and diagnostic subsets, the most reliably transferable component preserves SJ-associated polarity. Transferability alone therefore does not establish objective-correctness semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。