用语言驱动的主动推断框架,让通用人工智能从设计上更安全。
A Framework for Inherently Safer AGI through Language-Mediated Active Inference
- 用自然语言表示信念与偏好,实现透明可控的决策机制。
- 通过资源感知的自由能最小化,约束智能体理性范围。
- 模块化结构支持可组合的安全设计,适合追求本质安全的AI研发者。
本文提出一种结合主动推断与大语言模型的新框架,旨在从根源上构建安全的通用人工智能(AGI)。传统基于事后解释和奖励工程的安全方法存在根本局限。本框架将安全保障嵌入系统核心设计,通过自然语言媒介实现透明的信念表征与分层价值对齐。系统采用多智能体架构,智能体依据主动推断原则自组织,偏好与安全约束通过分层马尔可夫毯传递。具体机制包括:(1)在自然语言中显式分离信念与偏好;(2)通过资源感知的自由能最小化实现有界理性;(3)通过模块化智能体结构实现可组合的安全性。论文以抽象与推理基准(ARC)为核心,提出验证框架安全特性的实验路线。该方法提供了一条非事后添加、而是内生安全的AGI发展路径。
原文摘要 · Abstract (English)
This paper proposes a novel framework for developing safe Artificial General Intelligence (AGI) by combining Active Inference principles with Large Language Models (LLMs). We argue that traditional approaches to AI safety, focused on post-hoc interpretability and reward engineering, have fundamental limitations. We present an architecture where safety guarantees are integrated into the system's core design through transparent belief representations and hierarchical value alignment. Our framework leverages natural language as a medium for representing and manipulating beliefs, enabling direct human oversight while maintaining computational tractability. The architecture implements a multi-agent system where agents self-organize according to Active Inference principles, with preferences and safety constraints flowing through hierarchical Markov blankets. We outline specific mechanisms for ensuring safety, including: (1) explicit separation of beliefs and preferences in natural language, (2) bounded rationality through resource-aware free energy minimization, and (3) compositional safety through modular agent structures. The paper concludes with a research agenda centered on the Abstraction and Reasoning Corpus (ARC) benchmark, proposing experiments to validate our framework's safety properties. Our approach offers a path toward AGI development that is inherently safer, rather than retrofitted with safety measures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。