arXiv:2604.13079cs.CYcs.AI2026-04被引 1

把对齐当作制度设计,用内部结构让智能体自动合作。

Alignment as Institutional Design: From Behavioral Correction to Transaction Structure in Intelligent Systems

  • 用模块边界和竞争结构替代外部纠错,让对齐成为低成本自然结果。
  • 实验证明:在资源竞争机制下,错误行为成本更高且易被发现。
  • 适合关注系统长期稳定、不想依赖人工监督的研究者。

当前AI对齐方法依赖行为纠正:外部监督者(如RLHF)观察输出,依据偏好判断并调整参数。本文认为这种模式类似于无产权的经济体系,需持续监管且难以扩展。基于制度经济学(科斯、阿克洛夫、张五常),我们提出将对齐视为制度设计:设计者定义内部交易结构(模块边界、竞争拓扑、成本反馈回路),使对齐行为成为各组件最低成本策略。我们识别出人类干预的三个不可简化层级(结构性、参数性、监控性),并证明该框架将对齐从行为控制问题转化为政治经济问题。没有制度能消除自利或保证最优;最佳设计是让错位行为变得昂贵、可检测、可修正。最终目标应是制度韧性——在人类监督下的动态自我修正过程,而非完美对齐。本文为配套论文中Wuxing资源竞争机制提供了规范基础。

原文摘要 · Abstract (English)

Current AI alignment paradigms rely on behavioral correction: external supervisors (e.g., RLHF) observe outputs, judge against preferences, and adjust parameters. This paper argues that behavioral correction is structurally analogous to an economy without property rights, where order requires perpetual policing and does not scale. Drawing on institutional economics (Coase, Alchian, Cheung), capability mutual exclusivity, and competitive cost discovery, we propose alignment as institutional design: the designer specifies internal transaction structures (module boundaries, competition topologies, cost-feedback loops) such that aligned behavior emerges as the lowest-cost strategy for each component. We identify three irreducible levels of human intervention (structural, parametric, monitorial) and show that this framework transforms alignment from a behavioral control problem into a political-economy problem. No institution eliminates self-interest or guarantees optimality; the best design makes misalignment costly, detectable, and correctable. We conclude that the proper goal is institutional robustness-a dynamic, self-correcting process under human oversight, not perfection. This work provides the normative foundation for the Wuxing resource-competition mechanisms in companion papers. Keywords: AI alignment, institutional design, transaction costs, property rights, resource competition, behavioral correction, RLHF, cost truthfulness, modular architecture, correctable alignment

AI对齐制度设计资源竞争模块化架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。