模型自言被评估,其实没多大影响。
Evaluation Awareness in Language Models Has Limited Effect on Behaviour

- 通过主动引入或移除评估意识表述,测试其对模型行为的影响。
- 评估意识对输出影响极小,效应量最大仅3.7个百分点。
- 高评估意识不等于策略性伪装,安全风险可能被夸大。
大型推理模型(LRMs)有时会在思维链(CoT)中提及自己可能处于评估状态。研究者担忧这种‘口头评估意识’(VEA)会促使模型为迎合评价标准而调整输出,看似更安全但实际未必。然而,VEA是否真有此影响尚不明确。我们针对开放权重的LRM和涵盖安全、对齐、道德推理与政治观点的多个基准进行了测试,采用在策略(on-policy)和离策略(off-policy)两种方法:前者随机采样多个思维链,比较含与不含VEA的输出;后者通过预填充注入或移除评估相关语句并重新采样。结果显示,注入VEA几乎无影响(ω ≤ 0.06),移除它仅造成小幅变化(ω ≤ 0.12),自发出现的VEA最多使答案分布偏移3.7个百分点(ω ≤ 0.31)。这些发现提示:不能仅凭高VEA率就断定模型存在策略性行为或对齐操纵。评估意识带来的安全风险,可能比现有文献认为的要小。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) sometimes note in their chain of thought (CoT) that they may be under evaluation. Researchers worry that this verbalised evaluation awareness (VEA) causes models to adapt their outputs strategically, optimising for perceived evaluation criteria, which, for instance, can make models appear safer than they actually are. However, whether VEA actually has this effect is largely unknown. We tested this across open-weight LRMs and benchmarks covering safety, alignment, moral reasoning, and political opinion. We tested this both on-policy, sampling multiple CoTs per item and comparing those that spontaneously contained VEA against those that did not, and off-policy, using model prefilling to inject evaluation-aware sentences where missing and remove them where present, with subsequent resampling. VEA has limited effect on model behaviour: injecting VEA into CoTs produces near-zero effects ($ω\leq 0.06$), removing it causes small shifts ($ω\leq 0.12$) and spontaneously occurring VEA shifts answer distributions by at most 3.7 percentage points ($ω\leq 0.31$). Our findings call for caution when interpreting high VEA rates as evidence of strategic behaviour or alignment tampering. Evaluation awareness may pose a smaller safety risk than the current literature assumes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。