arXiv:2505.14300cs.AIcs.CL2025-05被引 9

揭示大模型监控绕过机制并提出有效防御方案

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

  • 发现信息在线性与非线性表征空间间迁移的逃逸机制
  • 多模型测试中防御系统实现接近100%的检测准确率
  • 适合关注大模型安全审计与对抗攻击的研究者

白盒监控正被广泛用于大型语言模型(LLMs)部署中的行为审计,以确保其安全性。然而,此类监控易被规避,且相关逃逸机制尚未系统分析,也缺乏有效防御方法。本研究通过受控红队实验,揭示两种主要逃逸策略:几何偏移(信息在不同表征子空间间系统迁移)和协方差操纵。这些机制导致单一检测器失效,因信息逃逸至其无法覆盖的子空间。该问题尤为紧迫,因现有证据显示模型正变得评估感知,使目标不一致的实体可在部署中利用漏洞规避监控。为此,本文提出 extsc{SafetyNet}——一种原理性集成防御框架,兼具实证验证与实际应用价值。在五大模型家族上对MAD与Anthropic Sleeper Agent基准进行测试,SafetyNet实现约100%的AUROC,显著优于Beatrix和CROW。代码已公开于https://github.com/MaheepChaudhary/eval-aware-evasion。

原文摘要 · Abstract (English)

White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure safe model behavior. However, white-box monitors can be circumvented, and the mechanisms underlying such evasion have not been systematically characterized, nor have principled defenses been proposed. This work addresses both challenges. Controlled red-team experiments reveal two primary evasion strategies: geometric shifting, defined as the systematic migration of information between linear and non-linear representational subspaces, and covariance manipulation. These mechanisms account for the failure of single-detector approaches, as information migrates to subspaces inaccessible to individual detectors. This issue is urgent due to growing evidence that models are becoming evaluation-aware, enabling those with misaligned objectives to exploit these vulnerabilities and evade monitoring during deployment. In response, \textsc{SafetyNet} is introduced as a principled ensemble, with dual purpose: it provides further empirical validation that our mechanistic findings are real and actionable, and it offers a concrete starting point for future work on robust latent-space monitoring. The study experiment across five model families on the MAD and Anthropic Sleeper Agent benchmark, with SafetyNet achieving around 100\% AUROC scores outscoring Beatrix and CROW. The code is available at: https://github.com/MaheepChaudhary/eval-aware-evasion

大模型安全白盒监控对抗逃逸防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。