通过组织路径监控AI能力增长,防止失控。
Mitigating loss of control in advanced AI systems through instrumental goal trajectories
- 提出三类组织路径追踪AI能力扩展轨迹
- 能力增长依赖算力/数据/资金等资源获取
- 适合关注AI治理与系统安全的研究者
高度智能的AI系统可能因追求工具性目标而削弱人类控制。现有缓解措施多聚焦技术层面:追踪系统能力、通过人类反馈强化学习调整行为,以及设计可纠正、可中断的系统。本文提出工具性目标轨迹(IGTs),将管控范围扩展至模型之外。能力提升通常依赖算力、存储、数据和邻近服务等技术资源,这些资源需通过组织内的采购、治理和金融三条路径获取。每条路径产生可监控的组织性成果,当系统能力或行为超出可接受阈值时,可作为干预点。IGTs为定义能力水平提供了具体途径,并拓展了可纠正性和可中断性的实现方式,促使关注重点从模型属性转向支撑其运行的组织系统。
原文摘要 · Abstract (English)
Researchers at artificial intelligence labs and universities are concerned that highly capable artificial intelligence (AI) systems may erode human control by pursuing instrumental goals. Existing mitigations remain largely technical and system-centric: tracking capability in advanced systems, shaping behaviour through methods such as reinforcement learning from human feedback, and designing systems to be corrigible and interruptible. Here we develop instrumental goal trajectories to expand these options beyond the model. Gaining capability typically depends on access to additional technical resources, such as compute, storage, data and adjacent services, which in turn requires access to monetary resources. In organisations, these resources can be obtained through three organisational pathways. We label these pathways the procurement, governance and finance instrumental goal trajectories (IGTs). Each IGT produces a trail of organisational artefacts that can be monitored and used as intervention points when a systems capabilities or behaviour exceed acceptable thresholds. In this way, IGTs offer concrete avenues for defining capability levels and for broadening how corrigibility and interruptibility are implemented, shifting attention from model properties alone to the organisational systems that enable them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。