arXiv:2608.16852cs.AI2026-08

现有合规检测器大多‘规则盲视’,无法真正理解法规内容。

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

论文配图:What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
图 1 · 摘自论文原文
  • 通过替换或删除规则,检测器准确率不变,暴露其依赖表面特征而非真实规则。
  • 设计跨规则交叉测试基准,验证了现有检测方法在复杂场景下失效。
  • 提出无需训练的内部合规得分(ICS),成本低但易被对抗攻击绕过。

部署语言模型的合规监控正日益作为法律与审计控制手段,检查模型输出是否符合数据保护、医疗、金融及平台政策等书面规则。此类监控仅在检测结果真正依赖于所列规则时才有效,而非依赖场景表面特征。我们发现当前合规检测器普遍存在‘规则盲视’问题:删除、打乱或替换核心规则后,所有测试的守门器和激活探针的检测准确率均无变化,包括一个能正确引用规则条款却在条款互换后仍维持原判的策略条件守门器。我们构建了一个专用于验证的基准,包含两个规则与两个场景,使单一规则无法预测标签,证实了该缺陷的存在;且只有逐步推理可规避此问题。由于需支持大规模无重训练审计,我们引入内部合规得分(ICS)——一种基于10个标注对、通过单个投影实现的无训练激活读出。我们将ICS置于与守门器相同的审查标准下检验:其未达到预注册的超越基线标准,且词袋模型恰好匹配其综合泛化能力。尽管如此,其低成本特性使其可用于审计四个已部署守门器、一个8B零样本裁判及十三个基准,提升响应排序的机械验证通过率,但自适应白盒攻击可消除此增益。我们公开反事实协议与交叉规则基准,以供未来探测与守门器声明的验证。

原文摘要 · Abstract (English)

Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart. A purpose-built benchmark crossing two rules with two scenarios, so that neither alone predicts the label, confirms the failure under a design no prior benchmark rules out, and shows that step by step reasoning, not any fast detector we test, is what escapes it. Auditing at scale requires a retraining-free detector, so we introduce the Internal Compliance Score (ICS): a training-free activation readout calibrated from ten labelled pairs and scored by a single projection. We hold ICS to the same scrutiny as the guards it audits: a pre-registered criterion for beating trivial baselines is not met, and a bag-of-words model matches its pooled generalisation exactly. It remains useful because it is inexpensive, letting us audit four deployed guard models, an 8B zero-shot judge, and thirteen benchmarks, and it raises the mechanically verified pass rate when used to rank candidate responses, though an adaptive white-box attack removes this gain. We release the counterfactual protocol and crossed-rule benchmark so rule blindness can be tested in future probe and guard claims.

合规检测规则盲视激活探针审计基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。