构建可审计的自治系统安全治理框架,实现决策与安全分离。
The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety
- 用提议者与安全探针解耦决策生成与安全控制
- 通过版本化更新安全探针修复多数新发安全问题
- 适合需要持续合规、可追溯的高风险自主系统
多智能体系统为角色分解、协调和规范治理提供了成熟方法,这些能力在日益强大的自主决策组件嵌入智能体系统时仍至关重要。尽管学习和生成模型显著扩展了系统能力,但其安全行为常与训练过程纠缠,导致透明度低、审计困难且部署后更新成本高。本文提出一种以治理为中心的混合多智能体架构——对齐飞轮(Alignment Flywheel),将决策生成与安全治理解耦。提议者(Proposer)生成候选轨迹,安全探针(Safety Oracle)通过稳定接口返回原始安全信号。执行层在运行时应用显式风险策略,治理多智能体系统则通过审计、不确定性驱动验证及版本化迭代监督探针。核心工程原则为补丁局部性:多数新观测到的安全失效可通过更新受控的探针及其发布流程解决,无需回滚或重训底层决策组件。该架构对提议者与安全探针的实现无特定要求,明确定义了角色、产物、协议与发布语义,支持运行时门控、审计输入、签名补丁及分布式环境中的分阶段上线。最终形成一套可集成高能力但不可靠自主系统的混合多智能体工程框架,实现显式、版本化、可审计的监督。
原文摘要 · Abstract (English)
Multi-agent systems provide mature methodologies for role decomposition, coordination, and normative governance, capabilities that remain essential as increasingly powerful autonomous decision components are embedded within agent-based systems. While learned and generative models substantially expand system capability, their safety behavior is often entangled with training, making it opaque, difficult to audit, and costly to update after deployment. This paper formalizes the Alignment Flywheel as a governance-centric hybrid MAS architecture that decouples decision generation from safety governance. A Proposer, representing any autonomous decision component, generates candidate trajectories, while a Safety Oracle returns raw safety signals through a stable interface. An enforcement layer applies explicit risk policy at runtime, and a governance MAS supervises the Oracle through auditing, uncertainty-driven verification, and versioned refinement. The central engineering principle is patch locality: many newly observed safety failures can be mitigated by updating the governed oracle artifact and its release pipeline rather than retracting or retraining the underlying decision component. The architecture is implementation-agnostic with respect to both the Proposer and the Safety Oracle, and specifies the roles, artifacts, protocols, and release semantics needed for runtime gating, audit intake, signed patching, and staged rollout across distributed deployments. The result is a hybrid MAS engineering framework for integrating highly capable but fallible autonomous systems under explicit, version-controlled, and auditable oversight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。