用强化学习训练更可信的AI审计员,能发现隐藏行为且误报极低。
Training Alignment Auditors via Reinforcement Learning

- 通过强化学习让AI审计员自主设计调查策略,提升审查深度。
- 在真实模型中检测到更多潜在危害行为,误报率低于1%。
- 可跨不同模型架构通用,适合高风险AI系统安全验证。
前沿模型的对齐审计越来越依赖大语言模型(LLM)审计员来大规模识别不良行为,但现有自动化审计员常因调查不连贯或缺乏现实性而表现不佳。本文提出使用强化学习改进LLM审计员:在最佳训练环境中,策略会针对可能包含由系统提示植入隐藏行为的目标模型进行调查。一个已知目标是否存在隐藏行为的LLM裁判,将策略生成的调查与参考调查整体对比以决定奖励。系统性消融实验表明,成对奖励比逐点奖励更鲁棒,且引入无植入行为的目标有助于保持低误报率。训练后,审计员在具有植入行为的目标上调查质量显著提升,对未修改生产模型中令人担忧行为的揭露率提高,同时审计真实性增强,误报率始终低于1%。此外,审计能力在不同模型架构间具备泛化性:在AuditBench的对抗微调目标上的表现大幅提升[Sheshadri et al., 2026]。
原文摘要 · Abstract (English)
Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. In this work, we improve LLM auditors with reinforcement learning. In our best training environment, the policy investigates target models that potentially possess hidden behaviors planted via their system prompt. An LLM judge, which knows whether the target has a hidden behavior, holistically compares the policy's investigation to a reference investigation to determine the reward. With systematic ablations, we find that pairwise rewards yield more robust training compared to pointwise rewards, and that adding targets without planted behaviors helps maintain a low false positive rate. Training improves investigation quality against targets with planted behaviors, the rate of concerning behaviors surfaced in unmodified production models, and audit realism, while false-positive rates stay below 1%. Furthermore, auditing capabilities generalize across scaffolds: performance on AuditBench's adversarially fine-tuned targets substantially improves [Sheshadri et al., 2026].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。