用可调保守性机制,让人类轻松掌控强于自己的智能体。
Calibrating Conservatism for Scalable Oversight

- 通过聚合多种评分函数,动态计算对高风险动作的惩罚
- 实测违规率精准匹配设定目标,且在有限时间内保证效果
- 适合需要可靠控制的强AI系统,如自主决策或伦理约束场景
具备自主规划与长期环境交互能力的智能体带来根本性的控制难题:人类如何保持对可能超越自身能力的系统的有效监督?现有可扩展监督方法依赖复杂假设、多为启发式,且缺乏适用于序列决策的统计保障。本文提出校准集体监督(CCO),将多种辅助评分函数聚合为偏离保守基线的惩罚项。受可实现效用保留启发,CCO实现集体保守性:动作惩罚与监督者关切程度成正比,只有当关切累积时才阻止高价值行动。CCO利用符合性决策理论在线校准保守性,确保不良结果低于用户指定阈值,具备有限时间边界且无需分布假设。在SWE-bench改进版上,弱监督者成功约束了对抗性错位的强智能体;在MACHIAVELLI上,CCO显著降低伦理违规,同时保持原始奖励。两组实验中,经验违规率均紧密贴合理论预测的目标值。
原文摘要 · Abstract (English)
Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain meaningful oversight of systems that may exceed their own capabilities? Existing approaches to scalable oversight rely on complex assumptions, remain largely heuristic, or lack practical methods for sequential settings with statistical guarantees. We introduce Calibrated Collective Oversight (CCO), which aggregates diverse auxiliary scoring functions into a penalty measuring deviation from a conservative baseline. Inspired by Attainable Utility Preservation, CCO enables collective conservatism: actions face a penalty proportional to overseer concern, so high-utility actions are still selected when overseers find them unobjectionable and overridden only when concern accumulates. CCO calibrates this conservatism online using Conformal Decision Theory, ensuring that undesirable outcomes remain below a user-specified target threshold with finite-time bounds and no distributional assumptions. On a modified version of SWE-bench, weaker overseers successfully constrain an adversarially misaligned stronger agent; on MACHIAVELLI, CCO substantially reduces ethical violations while preserving reward. In both settings, empirical violation rates closely match the specified targets, as predicted by the theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。