优化型AI无法真正响应规范,因其架构本质排斥道德边界与自主判断。
Agency and Architectural Limits: Why Optimization-Based Systems Cannot Be Norm-Responsive
- 提出代理能力的两个必要条件:不可协商的边界与非推理性暂停机制。
- 证明基于强化学习的LLM因优化本质,必然无法实现规范响应。
- 适合关注AI伦理、系统治理与人类监督风险的研究者阅读。
AI系统在医疗诊断、法律研究、金融分析等高风险场景中被广泛部署,常假设其可受规范约束。本文证明,基于优化的系统(特别是通过人类反馈强化学习训练的大语言模型)在形式上无法满足这一假设。真正的代理行为需同时具备两项必要且充分的架构条件:一是将某些边界视为不可协商的约束而非可权衡的权重(不可通约性);二是存在非推理性机制,在边界受威胁时能暂停处理(否定式响应)。基于强化学习的系统在本质上与这两项条件不相容。优化的核心操作——将所有价值统一到单一标量指标并始终选择得分最高输出——恰恰排除了规范治理与代理能力。这种不相容性并非可通过训练修复的技术缺陷,而是优化本身固有的形式约束。因此,已知的失败模式(谄媚、幻觉、非忠实推理)并非意外,而是结构性必然结果。当人类在指标压力下被迫验证AI输出时,会从真正代理人退化为仅执行标准检查的优化器,从而消除唯一能承担规范责任的主体。除理论证明外,本文的核心贡献是提出一种不依赖具体载体的架构规范,阐明任何系统(生物、人工或制度)要成为真正代理人所必须满足的基本条件。
原文摘要 · Abstract (English)
AI systems are increasingly deployed in high-stakes contexts (medical diagnosis, legal research, financial analysis) under the assumption they can be governed by norms. This paper demonstrates that the assumption is formally invalid for optimization-based systems, specifically Large Language Models trained via Reinforcement Learning from Human Feedback (RLHF). Genuine agency requires two necessary and jointly sufficient architectural conditions. First, the capacity to maintain certain boundaries as non-negotiable constraints rather than tradeable weights (Incommensurability). Second, a non-inferential mechanism capable of suspending processing when those boundaries are threatened (Apophatic Responsiveness). RLHF-based systems are constitutively incompatible with both conditions. The operations that make optimization powerful, unifying all values on a scalar metric and always selecting the highest-scoring output, are precisely the operations that preclude normative governance and agency. This incompatibility is not a correctable training bug awaiting a technical fix. It is a formal constraint inherent to what optimization is. Consequently, documented failure modes (sycophancy, hallucination, and unfaithful reasoning) are not accidents but expected structural manifestations. Misaligned deployment triggers a second-order risk termed the Convergence Crisis. When humans are forced to verify AI outputs under metric pressure, they degrade from genuine agents into criteria-checking optimizers, eliminating the only component capable of bearing normative accountability. Beyond the incompatibility proof, this paper's primary positive contribution is a substrate-neutral architectural specification deriving what any system (biological, artificial, or institutional) must necessarily satisfy to qualify as a genuine agent rather than a sophisticated instrument.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。