arXiv:2609.03026cs.LGcs.AI2026-09

提出测试模型内部评估器是否适合干预与控制的基准框架。

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

  • 构建固定任务合约,分离评估精度与行动导致的损失。
  • 实验证明准确评估不等于选对动作,需同时关注两者。
  • 适用于安全监控、电路干预等场景,适合可解释性研究者使用。

机制可解释性被广泛用于指导激活调节、电路移除和安全监控等干预行为。然而,即使内部估计平均准确,仍可能选择较差的行动。本文提出 ObserverBench,一个用于测试内部估计器(观察者)在干预、控制或安全任务中是否足够的基准框架。每个任务固定模型、信息边界、允许动作、决策规则、未见样本及损失函数。该基准分别报告估计准确性与所选行动造成的损失。理论与实验表明两者缺一不可。在闭环控制中,观察者误差在起始点及干预可达方向上至关重要。在 GPT-2-small 与 Qwen2.5-7B 的电路干预任务中,成对观察者能更准确预测未见效应,但未必选择更优行动;基于行动损失训练的观察者则选择更低损失行动。在安全筛查中,完美区分违规的评分,在违规成本不同时可能导致干预预算分配不当。在 Qwen2.5-7B、Gemma-2-9B-it 及前瞻性冻结的 Qwen3.5-9B APPS 任务中,AUROC 排名监测器顺序与部署损失不同,最佳信息源随模型变化。稀疏 SAE 读出在报告的 Qwen 面板上,因激活密度或检查点不匹配,仍落后于层匹配的密集控制。ObserverBench 提供固定任务合约、可运行基线与表格提交方式,通过行动效果评估可解释性方法。

原文摘要 · Abstract (English)

Mechanistic interpretability is increasingly used to guide interventions such as activation steering, circuit removal, and safety monitoring. Yet an internal estimate that is accurate on average can still choose a poor action. We present ObserverBench, a benchmark framework for testing whether an internal estimator---an observer---is adequate for the intervention, control, or safety task it directs. Each task fixes the model, information boundary, allowed actions, decision rule, held-out cases, and loss. The benchmark reports estimation accuracy separately from the loss caused by the chosen action. Theory and experiments show why both are needed. In closed-loop control, observer errors matter at the starting point and along directions the allowed intervention can reach. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers predict unseen effects more accurately without always choosing better actions; observers trained on action loss choose lower-loss actions. In safety triage, a score that perfectly separates violations can allocate a fixed intervention budget poorly when violations have different costs. Across Qwen2.5-7B, Gemma-2-9B-it, and prospectively frozen Qwen3.5-9B APPS tasks, AUROC can rank monitors differently from deployment loss, and the best information source changes across models. Sparse SAE readouts also trail their layer-matched dense controls on the reported Qwen panels, under disclosed activation-density or checkpoint mismatches. ObserverBench provides fixed task contracts, runnable baselines, and table-based submissions for evaluating interpretability methods through the actions they enable.

可解释性干预评估基准测试模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。