让AI通过软最大化人类长期权力来实现安全与福祉的平衡
Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power
- 设计可调节的客观函数,量化人类长期权力并兼顾多元目标
- 基于世界模型用反向归纳或多智能体强化学习计算该指标
- 适合关注AI安全与人类权力平衡的研究者和开发者
权力是人工智能安全中的核心概念:权力追求作为工具性目标、人类权力的突然或渐进削弱、人机交互及国际AI治理中的权力均衡。同时,权力作为实现多样化目标的能力,对福祉至关重要。本文探讨通过显式促使AI赋能人类,并以理想方式管理人机间权力平衡,来兼顾安全与福祉。采用部分公理化的方法,设计了一个可参数化且可分解的目标函数,代表一种对不平等和风险均持谨慎态度的长期人类权力聚合度量。该函数考虑了人类认知局限与社会规范,并关键地涵盖多种可能的人类目标。我们推导出基于给定世界模型,通过反向归纳或某种形式的多智能体强化学习近似计算该度量的算法。在多种典型情境中演示了(软性)最大化此度量的后果,并描述其可能引发的工具性子目标。我们的审慎评估认为,软性最大化合适的人类权力聚合度量,可能成为比直接效用目标更安全的代理型AI系统有益目标。
原文摘要 · Abstract (English)
Power is a key concept in AI safety: power-seeking as an instrumental goal, sudden or gradual disempowerment of humans, power balance in human-AI interaction and international AI governance. At the same time, power as the ability to pursue diverse goals is essential for wellbeing. This paper explores the idea of promoting both safety and wellbeing by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach, we design a parametrizable and decomposable objective function that represents an inequality- and risk-averse long-term aggregate of human power. It takes into account humans' bounded rationality and social norms, and, crucially, considers a wide variety of possible human goals. We derive algorithms for computing that metric by backward induction or approximating it via a form of multi-agent reinforcement learning from a given world model. We exemplify the consequences of (softly) maximizing this metric in a variety of paradigmatic situations and describe what instrumental sub-goals it will likely imply. Our cautious assessment is that softly maximizing suitable aggregate metrics of human power might constitute a beneficial objective for agentic AI systems that is safer than direct utility-based objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。