设计一种让AI主动赋能人类、平衡人机权力的新型目标函数。
A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

- 基于人类理性局限与社会规范,构建可调节的人类权力长期聚合目标函数。
- 能适应多样人类目标,避免权力失衡,提升人机互动安全性与公平性。
- 适用于注重伦理与长期安全的AI系统设计,适合研究人机共生的学者。
本文提出一种基于理想特性(desiderata)的原理性方法,旨在通过显式约束人工智能代理来促进人类福祉与安全,确保人机之间的权力关系处于理想状态。设计了一个可参数化且可分解的目标函数,用于衡量不平等和风险厌恶下的长期人类权力累积。该函数能融合人类有限理性和社会规范模型,并涵盖多种可能的人类目标。论文证明了特定理想特性如何强制特定函数形式并限制参数范围。通过在若干典型场景中软最大化该指标,分析其可能引发的行为后果及潜在工具性子目标。
原文摘要 · Abstract (English)
This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach based on desirable properties, we design a parametrizable and decomposable objective function for AI systems that represents an inequality- and risk-averse long-term aggregate of human power. It can take into account models of human bounded rationality and social norms, and crucially, considers a wide variety of possible human goals. We prove how certain desiderata enforce particular functional forms and restrict parameter ranges. We exemplify the consequences of softly maximizing this metric in several paradigmatic situations and describe what instrumental sub-goals it will likely imply.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。