通过约束强化学习,让高风险智能体在授权范围内动态调整决策权。
Selection as Power: Constrained Reinforcement for Bounded Decision Authority
- 用外部主权约束下的强化更新控制选择权力,防止过度集中。
- 无约束时高学习率下决策会快速固化为单一模式,而有约束则能持续优化。
- 适合关注智能体治理与安全的系统设计者和研究人员。
选择即权力理论指出,上游选择权而非内部目标错配,是高风险智能体系统的主要风险源。然而原框架为静态:治理约束虽限制选择权但不随时间调整。本文将其扩展至动态场景,引入激励型选择治理机制,即在外部强制主权约束下对评分与筛选参数进行强化更新。我们形式化了选择为受限强化过程,其中参数更新被投影到治理定义的可行集上,防止超出规定范围的集中。在多个受监管金融场景中,无约束强化在重复反馈下始终趋于确定性主导,尤其在高学习率时;而激励治理则实现自适应改进且保持选择集中度受控。基于投影的约束将强化从不可逆锁定转变为可控适应,治理债务量化了优化压力与权限边界之间的张力。结果表明,当主权约束在每一步更新中被强制执行时,学习动态可与结构多样性共存,为高风险智能体系统集成强化学习提供了原则性路径,同时不放弃有限的选择权。
原文摘要 · Abstract (English)
Selection as Power argued that upstream selection authority, rather than internal objective misalignment, constitutes a primary source of risk in high-stakes agentic systems. However, the original framework was static: governance constraints bounded selection power but did not adapt over time. In this work, we extend the framework to dynamic settings by introducing incentivized selection governance, where reinforcement updates are applied to scoring and reducer parameters under externally enforced sovereignty constraints. We formalize selection as a constrained reinforcement process in which parameter updates are projected onto governance-defined feasible sets, preventing concentration beyond prescribed bounds. Across multiple regulated financial scenarios, unconstrained reinforcement consistently collapses into deterministic dominance under repeated feedback, especially at higher learning rates. In contrast, incentivized governance enables adaptive improvement while maintaining bounded selection concentration. Projection-based constraints transform reinforcement from irreversible lock-in into controlled adaptation, with governance debt quantifying the tension between optimization pressure and authority bounds. These results demonstrate that learning dynamics can coexist with structural diversity when sovereignty constraints are enforced at every update step, offering a principled approach to integrating reinforcement into high-stakes agentic systems without surrendering bounded selection authority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。