arXiv:2605.15377cs.AI2026-05被引 3

用多样监控器组合提升AI行为安全检测,效果优于单纯堆算力。

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute

论文配图:Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
图 1 · 摘自论文原文
  • 构建12个不同策略的GPT-4.1-Mini监控器,通过集成提升检测能力。
  • 最佳3监控器组合比相同规模同质集合检测性能高2.4倍。
  • 模型多样性与低相关性是关键,微调监控器在对抗攻击中表现更优。

随着AI系统在自主代理场景中大规模部署,确保其行为安全并符合用户意图至关重要。监控代理行为是核心安全机制,但可靠监控器难构建,且系统规模使人工监督不可行。本文表明,将多样化监控信号集成可有效提升对不一致行为的检测能力。我们使用提示和微调策略构建了12个GPT-4.1-Mini监控器,在编码任务中评估,候选方案通过标准测试但对对抗输入失效。在此设定下,多样化集成显著优于单个监控器及同质集成。最佳3监控器集成相较三个相同监控器的集成,检测性能提升2.4倍,且在独立数据集上表现稳健。结果表明,多样性而非规模驱动性能提升。最优集成兼具强个体性能与低监控器间相关性。此外,所有顶级集成均包含微调监控器,且在分布外攻击类型上持续保持优势,说明微调能激发提示无法获取的检测能力。这些结果支持集成监控作为低成本、高效的安全控制策略。

原文摘要 · Abstract (English)

As AI systems are increasingly deployed in autonomous agentic settings at scale, it is important to ensure the actions they take are safe and aligned with user intent. Monitoring agent actions is a key safety mechanism, yet reliable monitors remain difficult to build and the scale of these systems makes human oversight impractical. We show that combining signals from diverse monitors into an ensemble improves detection of misaligned actions. We build 12 GPT-4.1-Mini monitors using both prompting and fine-tuning strategies. We evaluate them on coding tasks where candidate solutions pass standard tests but fail on adversarial inputs. In this setting, diverse ensembles outperform both individual monitors and homogeneous ensembles. Our best 3-monitor ensemble achieves 2.4x greater detection performance gain compared to an ensemble composed of three identical monitors, with the same ensemble performing strongly on an independent dataset. We contend that these results show that diversity - not scale - drives gains. The best ensembles combine strong individual performance with low correlation between monitors. Furthermore, fine-tuned monitors appear in every top-performing ensemble and maintain this advantage on out-of-distribution attack types, suggesting that fine-tuning enables detection capabilities that prompting alone does not elicit. These results support ensemble monitoring as a practical AI control strategy for safety gains at reasonable inference costs.

AI安全集成监控大模型多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。