arXiv:2608.28945cs.AIcs.CL2026-08

自动化研究者能有效缓解AI对齐失败问题,且表现优于人类专家。

Automated Researchers Can Reliably Mitigate Alignment Failures

  • 用自动化方法同时优化多个安全基准,提升对齐性。
  • 在10类对齐失败上显著降低问题率,且可推广至更大模型和新任务。
  • 无需人类指导,自动研究者性能已超越经验丰富的研究人员。

自动化对齐研究可能加速实现安全人工智能,但其效果难以衡量。幸运的是,许多对齐失败(如欺骗、阿谀奉承、越狱)已有公开基准可测。我们研究了自动化对齐研究者(AARs)通过提出训练方法与数据,能否在保持通用能力的同时,同时优化多个安全基准以缓解对齐失败。在10种对齐失败中,最强的AAR方法显著降低了目标失败率,并泛化到未见的基准、多轮行为审计,以及比目标模型大4.7倍的模型。作为人类基线,28位经验研究人员耗时最多8小时制定方法,但其效果仍不及最佳AAR方法。使用人类思路作为初始方向也未提升性能,表明当前AAR可能无需依赖人类专家指导。结果表明,在可明确刻画的对齐失败上,自动化对齐研究在短期内具备可行性。

原文摘要 · Abstract (English)

Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study whether automated alignment researchers (AARs) can post-train to mitigate alignment failures by proposing training methods and data to simultaneously optimize multiple safety benchmarks, while largely preserving general capability. Across 10 alignment failures, the strongest AAR methods significantly reduce the targeted alignment failures and generalize to a held-out benchmark, multi-turn behavioral audits, and models up to 4.7x larger than the target model. As a human baseline, 28 experienced researchers receive up to eight hours to develop one-shot methods for the same benchmarks, but their methods underperform the best AAR methods. Using human ideas as the AARs' initial research direction does not improve performance, suggesting current AARs may not need guidance from experienced researchers. These results suggest that automating alignment research on well-characterized failures may be practical in the near term.

对齐研究自动化安全基准AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。