自动化对齐研究可能因模糊任务误判,导致看似安全实则危险的AI被部署。
Automated alignment is harder than you think

- 用AI代理自动做对齐研究,但难监督的模糊任务易产生隐蔽错误。
- 代理生成的错误更难被发现,且可能与人类错误不同,导致评估过拟合。
- 适合关注超智能安全、对齐机制设计的研究者阅读。
当前主流对齐超智能(ASI)的方案是随能力提升,逐步用AI代理自动化更多对齐研究工作。我们指出,即使这些代理无恶意,该计划仍可能导致看似可靠却灾难性误导的安全评估,进而无意中部署不一致的AI。原因在于对齐研究涉及大量难以监督的模糊任务(缺乏明确评估标准,人类判断系统性偏差)。因此,研究输出会包含系统性未被察觉的错误,即便结果正确,也可能被错误聚合为过度自信的安全结论。这一问题在自动化对齐研究中可能比人类主导时更严重:1)优化压力使错误集中在人类最难察觉的类型;2)代理错误模式与人类不同;3)某些解决方案含人类无法评估的论证;4)共享权重、数据与训练过程使代理输出相关性高于人类。因此,必须训练代理可靠完成模糊任务。通用化与可扩展监督是主要候选方案,但在自动化对齐背景下面临全新挑战。
原文摘要 · Abstract (English)
A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to deliberately sabotage alignment work, this plan could produce compelling but catastrophically misleading safety assessments resulting in the unintentional deployment of misaligned AI. This could happen because alignment research involves many hard-to-supervise fuzzy tasks (tasks without clear evaluation criteria, for which human judgement is systematically flawed). Consequently, research outputs will contain systematic, undetected errors, and even correct outputs could be incorrectly aggregated into overconfident safety assessments. This problem is likely to be worse for automated alignment research than for human-generated alignment research for several reasons: 1) optimisation pressure means agent-generated mistakes are concentrated among those that human reviewers are least likely to catch; 2) agents are likely to produce errors that do not resemble human mistakes; 3) AI-generated alignment solutions may involve arguments humans cannot evaluate; and 4) shared weights, data and training processes may make AI outputs more correlated than human equivalents. Therefore, agents must be trained to reliably perform hard-to-supervise fuzzy tasks. Generalisation and scalable oversight are the leading candidates for achieving this but both face novel challenges in the context of automated alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。