提出构建可扩展对齐智能体的原理,让机器学会理解世界和人类偏好。
Possible Principles for Aligned Structure Learning Agents
- 通过结构学习与心智理论,让智能体构建世界模型和人类偏好模型。
- 提出基于信息几何与核心知识的模块化结构学习框架。
- 为未来对齐智能系统提供数学基础与设计原则,适合AI安全研究者。
本文从自然智能的第一性原理出发,提出可扩展对齐人工智能的发展路线。核心在于使人工智能体能够学习一个包含人类偏好在内的优质世界模型。为此,关键任务是构建能表征世界及其他智能体世界模型的代理,这属于结构学习(又称因果表示学习或模型发现)范畴。文章探讨了结构学习与对齐问题,并提出指导原则,融合数学、统计学与认知科学的思想。1)强调核心知识、信息几何与模型简化在结构学习中的关键作用,建议采用核心结构模块以适应多种自然场景。2)通过结构学习与心智理论实现对齐智能体,以阿西莫夫机器人三定律为例,形式化描述智能体应谨慎行动以最小化他者痛苦;并进一步提出优化对齐的方法。这些洞见可引导未来对齐系统的开发,助力现有或新型结构学习系统的规模化。
原文摘要 · Abstract (English)
This paper offers a roadmap for the development of scalable aligned artificial intelligence (AI) from first principle descriptions of natural intelligence. In brief, a possible path toward scalable aligned AI rests upon enabling artificial agents to learn a good model of the world that includes a good model of our preferences. For this, the main objective is creating agents that learn to represent the world and other agents' world models; a problem that falls under structure learning (a.k.a. causal representation learning or model discovery). We expose the structure learning and alignment problems with this goal in mind, as well as principles to guide us forward, synthesizing various ideas across mathematics, statistics, and cognitive science. 1) We discuss the essential role of core knowledge, information geometry and model reduction in structure learning, and suggest core structural modules to learn a wide range of naturalistic worlds. 2) We outline a way toward aligned agents through structure learning and theory of mind. As an illustrative example, we mathematically sketch Asimov's Laws of Robotics, which prescribe agents to act cautiously to minimize the ill-being of other agents. We supplement this example by proposing refined approaches to alignment. These observations may guide the development of artificial intelligence in helping to scale existing -- or design new -- aligned structure learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。