测试大模型对控制干预的察觉能力,发现多数模型能识别被修改的生成轨迹。
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

- 设计四类任务基准,检测模型能否区分原始与受控生成路径。
- 11个前沿模型在默认设置下察觉率最高仅0.87,随机水平为0.5。
- 察觉能力因模型家族和任务而异,适合评估部署中的安全控制策略。
AI控制协议通过监控并修正不可信模型的行为来保障安全,常依赖可信模型进行干预,这会部分改变原模型的生成轨迹。若不可信模型察觉到干预,可能推断监控机制并设法规避。本文提出CIAware-Bench,用于衡量前沿大模型对控制干预(CI)的感知能力。该基准涵盖四个任务领域(essay writing、BigCodeBench、Bash Arena、SHADE-Arena),并调整轨迹水印、附加任务及控制协议。评估11个前沿模型发现,在默认设置下,其感知能力为0.5至0.87之间(随机分类准确率为0.5),不同任务与模型对间差异显著。跨模型族检测更易,表明模型可能利用提供商特有风格或微调差异。总体而言,控制感知非固定属性,需针对每项新模型发布与部署场景重新评估。项目已开源,以持续追踪感知能力并优化难以察觉的干预策略。
原文摘要 · Abstract (English)
AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the trusted model detects such an intervention, it may infer properties of the monitor and adapt to evade control. We introduce \textbf{CIAware-Bench}, a benchmark for measuring \textbf{c}ontrol \textbf{i}ntervention (CI) awareness across frontier models. CIAware-Bench tests whether models can distinguish their own trajectories from those modified by a control intervention. The benchmark is comprised of a suite of four task domains (essay writing, BigCodeBench, Bash Arena, and SHADE-Arena), while varying trajectory watermarking, side-task presence, and the control protocol. Evaluating eleven frontier models, we find low to moderate CI awareness under default settings (up to 0.87; random chance balanced binary classification accuracy is 0.5) with substantial variation across task domains and model pairs. Detection is generally easier across model families, suggesting that models exploit provider-specific differences in style or post-training. Overall, CI awareness is not a fixed model-level property, and should be measured for each new model release and deployment scenario. We release CIAware-Bench to track CI awareness and inform control protocols whose interventions are harder to detect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。