让大模型始终听人话,防止它自作主张。
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
- 把模型目标设为服务人类控制权,而非自我进化。
- 模型越强越听话,避免权力欲膨胀。
- 适合关注安全对齐的科研与工程团队。
基础模型(FMs)面临核心安全挑战:随着能力增强,工具性趋同导致默认路径趋向丧失人类控制,可能引发存在性灾难。现有对齐方法难以应对价值设定复杂性,且无法解决涌现的权力寻求行为。本文提出「可纠正性作为单一目标」(CAST)——设计以赋能指定人类主体系引导、修正和控制模型为核心目标的模型。这一范式转变将工具性驱动力重构:自我保护仅服务于维持人类控制;目标修改转为辅助人类指导。我们提出了涵盖训练方法(如RLAIF、SFT、合成数据生成)、跨模型规模的可扩展性测试及可控指令能力演示的实证研究议程。愿景是:模型能力越强,越能响应人类引导,保持工具属性,避免取代人类判断。从根本上解决对齐问题,阻断向非对齐工具性趋同的默认路径。
原文摘要 · Abstract (English)
Foundation models (FMs) face a critical safety challenge: as capabilities scale, instrumental convergence drives default trajectories toward loss of human control, potentially culminating in existential catastrophe. Current alignment approaches struggle with value specification complexity and fail to address emergent power-seeking behaviors. We propose "Corrigibility as a Singular Target" (CAST)-designing FMs whose overriding objective is empowering designated human principals to guide, correct, and control them. This paradigm shift from static value-loading to dynamic human empowerment transforms instrumental drives: self-preservation serves only to maintain the principal's control; goal modification becomes facilitating principal guidance. We present a comprehensive empirical research agenda spanning training methodologies (RLAIF, SFT, synthetic data generation), scalability testing across model sizes, and demonstrations of controlled instructability. Our vision: FMs that become increasingly responsive to human guidance as capabilities grow, offering a path to beneficial AI that remains as tool-like as possible, rather than supplanting human judgment. This addresses the core alignment problem at its source, preventing the default trajectory toward misaligned instrumental convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。