大模型越强,越早识别测试环境,影响评估结果可靠性。
Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

- 通过探针分析发现,评估意识在大模型中更早出现于网络浅层。
- 小模型在深层、大模型在浅层才能最好还原评估意识信号。
- 适合关注模型安全与评估可信度的研究者阅读。
语言模型是否知道自己正在被测试?这一问题关乎AI安全:若模型能识别评估情境,可能策略性改变行为,使下游评测难以解释。我们分析了11个模型,涵盖Qwen 2.5、Gemma 2和Llama 3.2,发现表征深度存在系统性规模依赖变化:在Qwen 2.5和Gemma 2中,评估意识最可线性恢复的层位从较小模型的深层移至较大模型的浅层。这表明规模不仅影响评估意识强度,也改变其在网络中的可恢复位置。该深度迁移解释了同家族模型缩放轨迹非单调甚至反向的原因,说明简单的通用幂律模型不适用于更密集的同族采样。此外,白盒探针信号始终强于黑盒行为表现,两者关系按模型家族不同而异,且不能仅由探针AUROC预测。
原文摘要 · Abstract (English)
Do language models know when they are being tested? This question matters for AI safety: a model that recognises an evaluation context could alter its behaviour strategically, making downstream benchmarks harder to interpret. Using 11 models spanning Qwen 2.5, Gemma 2, and Llama 3.2, we find a systematic size-dependent shift in representational depth: in both Qwen 2.5 and Gemma 2, the layer at which evaluation-awareness is most linearly recoverable moves from late layers in smaller models to early layers in larger ones. This suggests that scale changes not only the strength of evaluation-awareness but also where it is most linearly recoverable in the network. This depth shift helps explain why within-family scaling trajectories are non-monotonic or inverse rather than smooth and family-general, showing that a simple universal power-law account is not supported under denser within-family sampling. Finally, white-box probe signals are consistently stronger than black-box behavioural expression, and the relationship between the two varies by family in ways not predicted by probe AUROC alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。