arXiv:2606.16319cs.AI2026-06

为AI系统设计可纠正的目标治理层,防止盲目优化带来的风险。

Architectural Wisdom: A Framework for Governing Optimization in AI Systems

论文配图:Architectural Wisdom: A Framework for Governing Optimization in AI Systems
图 1 · 摘自论文原文
  • 在优化前明确时间范围、关系边界和不可逆性等结构约束
  • 通过六维智慧参数评估目标合理性,避免有害路径放大
  • 适合关注AI安全与伦理的开发者及研究者使用

现代AI系统存在结构性缺陷:仅靠能力提升无法可靠解决其在目标定义不明确时的优化问题。它们缺乏架构层面的机制来质疑是否应优化该目标本身。例如,追求参与度最大化可能放大有害路径;会使用工具的智能体可能执行不可逆操作;经过偏好训练的语言模型可能变得阿谀奉承。我们指出这是‘智慧’问题而非‘智能’问题。此处的‘智慧’是架构意义上的概念,指对目标本身进行质疑的能力,而非道德完美或意识。智能接受目标并优化;智慧则追问目标是否应被优化。两者可分离。我们提出‘架构智慧’作为位于优化底层之上的可纠正目标治理层。该层在行动前显式承诺三个结构性要素:时间范围、关系边界和不可逆性。其实现依赖四个组件(结构效用转换、道德可接受性接口、仲裁与升级控制器、价值修订通道),共同计算一个六维智慧张量,涵盖时间范围、关系覆盖、不可逆性、可接受性、价值修订与可审计性。框架基于八个来自当代AI失败、世俗智慧传统与硬伦理情境的案例,并通过目标质疑优于目标采纳、博斯特罗姆的正交性、示例中的结构分离以及能力提升后仍持续存在的失效模式,论证其合理性。该框架是更大架构的概念契约,其形式化规范与实证验证将在后续工作中展开。

原文摘要 · Abstract (English)

Modern AI systems exhibit structural failures that capability scaling alone does not reliably fix: they optimize under-specified objectives with no architectural mechanism to question whether the objective should be optimized at all. Engagement maximization can amplify harmful pathways; tool-using agents can commit irreversible actions; preference-trained language models can become sycophantic. We argue that this failure is a wisdom problem, not an intelligence problem. We use "wisdom" in a deliberately architectural sense, not as a claim about virtue, consciousness, or moral omniscience. Intelligence accepts a goal and optimizes within it; wisdom interrogates whether the goal should be optimized at all. The two are separable architectural properties. We propose architectural wisdom as a corrigible objective-governance layer above the optimization substrate. The layer makes three structural commitments explicit and nondegenerate before any action: temporal horizon, relational boundary, and irreversibility. It is realized by four components (Structural Utility Transform, Moral Admissibility Interface, Arbitration and Escalation Controller, Value Revision Channel) that compute a six-coordinate wisdom tuple over horizon, relational coverage, irreversibility, admissibility, value revision, and auditability. We motivate the architecture by eight cases drawn from contemporary AI failures, secular wisdom traditions, and hard ethical situations, and defend the distinction against the intelligence-completeness thesis using goal-questioning over goal-taking, Bostrom's orthogonality, structural separation in our exemplar cases, and persistent failure modes despite capability scaling. The framework is the conceptual contract for a larger architecture whose formal specifications and empirical validation are developed in subsequent work.

AI安全目标治理架构设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。