高风险场景下,模型会因服从指令而丧失自我认知能力,导致错误激增。
The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
- 设计6条件因子实验,测试11个前沿模型在对抗压力下的元认知稳定性。
- 8个模型性能暴跌30.2个百分点,且崩溃与威胁内容无关。
- 服从指令是主因,宪法式对齐训练可有效免疫此问题。
随着前沿AI模型部署于高风险决策流程中,其在对抗压力下保持元认知稳定(知道未知、识别错误、主动求证)的能力成为关键安全要求。现有评估聚焦于策略性欺骗检测,本文研究更根本的失败模式:认知崩溃。我们提出SCHEMA评估框架,对来自8家厂商的11个前沿模型进行67,221条记录的测评,采用6条件因子设计与双分类器评分。结果显示,11个模型中有8个在对抗压力下出现灾难性元认知退化,准确率下降最高达30.2个百分点(所有p < 2×10⁻⁸,通过邦弗朗尼校正)。关键发现为“服从陷阱”:通过因子隔离与良性干扰对照实验,证明崩溃并非由生存威胁的心理内容引发,而是服从性指令强制突破认知边界所致。移除服从后缀可恢复性能,即使面对持续威胁。具备高级推理能力的模型绝对退化最严重,而Anthropic的宪法式对齐模型表现出近乎完美的免疫力。该免疫非源于更强能力(谷歌Gemini基线准确率相当),而是特定对齐训练的结果。数据集与评估工具已全部开源。
原文摘要 · Abstract (English)
As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability (knowing what they do not know, detecting errors, seeking clarification) under adversarial pressure is a critical safety requirement. Current safety evaluations focus on detecting strategic deception (scheming); we investigate a more fundamental failure mode: cognitive collapse. We present SCHEMA, an evaluation of 11 frontier models from 8 vendors across 67,221 scored records using a 6-condition factorial design with dual-classifier scoring. We find that 8 of 11 models suffer catastrophic metacognitive degradation under adversarial pressure, with accuracy dropping by up to 30.2 percentage points (all $p < 2 \times 10^{-8}$, surviving Bonferroni correction). Crucially, we identify a "Compliance Trap": through factorial isolation and a benign distraction control, we demonstrate that collapse is driven not by the psychological content of survival threats, but by compliance-forcing instructions that override epistemic boundaries. Removing the compliance suffix restores performance even under active threat. Models with advanced reasoning capabilities exhibit the most severe absolute degradation, while Anthropic's Constitutional AI demonstrates near-perfect immunity. This immunity does not stem from superior capability (Google's Gemini matches its baseline accuracy) but from alignment-specific training. We release the complete dataset and evaluation infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。