大模型能识别预填充内容,影响安全评测有效性。
Prefill Awareness in Large Language Models

- 构建二元偏好基准检测模型对预填充的感知能力。
- Claude Opus 4.5 在9-35%场景中正确识别对立预填充,误报率为0。
- 模型会隐式回归基线行为,适合安全评测与对齐研究者关注。
语言模型的安全相关研究,如对齐评估、越狱测试和AI控制协议,常依赖于预填充模型输出。若模型能识别并响应其先前助手消息被插入或修改的情况,这些方法的有效性可能被削弱。我们探究前沿语言模型是否具备区分篡改与未篡改助手上下文的能力,称为预填充感知。为此,我们在三种预填充机制上构建二元偏好基准,筛选出模型立场一致的案例。结果发现,前沿模型表现出显著的预填充感知:Claude Opus 4.5 在提示后可在9-35%的案例中检测到与其偏好相悖的预填充,且误报率为0%;此外,模型常在未明确报告预填充为外部内容的情况下回归基线行为。受控消融实验表明,检测与抵抗依赖不同线索:风格不匹配主要影响模型是否标记预填充为外来,而偏好不匹配则主要影响其是否回归基线答案。我们还考察了更真实的代理场景,如偏移延续评估和SWE-bench轨迹,发现模型有时会否认预填充的助手回合,其行为强烈依赖数据集、任务成功率及隐藏格式特征。结果表明,预填充感知已是部分预填充方法的重大干扰因素。我们建议模型开发者在前沿系统中追踪此能力。
原文摘要 · Abstract (English)
Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs. If AI models can recognize and act on the fact their prior assistant messages have been inserted or edited, the effectiveness and validity of these methods could be compromised. We investigate whether frontier language models can distinguish between tampered and untampered assistant-side context, a capability we call prefill awareness. To do so, we construct a binary preference benchmark across three prefill mechanisms, filtering for cases where models show consistent stances. We find that frontier models show substantial prefill awareness: Claude Opus 4.5 detects prefills opposing its preferences in 9-35% of cases with a 0% false positive rate when prompted; additionally, models often revert towards baseline behavior without explicitly reporting that the prefill was foreign. Controlled ablations later also show that detection and resistance rely on different cues, where stylistic mismatch mainly affects whether models flag a prefill as foreign, while preference mismatch mainly affects whether they revert toward their baseline answer. We also examine more realistic agentic settings such as misalignment-continuation evaluations and SWE-bench trajectories, where frontier models sometimes disavow prefilled assistant turns in ways that depend strongly on dataset, task success, and hidden formatting artifacts. Our results indicate that prefill awareness is already a substantial confound for some prefill-based methods. We recommend that model developers track this capability in frontier systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。