用多重验证机制防范AI生成的危险操作,提升云系统安全性。
Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructure
- 将操作提案转化为带证据链的声明式合同,交由多个隔离验证者评审。
- 在500个场景中,误批准危险操作率从18.5%降至0.3%。
- 适合需要高安全性的云原生AI操作管控场景。
随着大语言模型代理被集成到自治云运维中,分布式系统面临语义可靠性问题:提议代理可生成语法正确且静态授权的生产变更(如修改IAM策略、开放防火墙安全组或执行数据导出),但这些操作在实际运行中不安全。传统分布式共识协议仅复制确定性状态转换,不评估提议意图的安全性。为此,我们提出语义共识保证(Semantic Quorum Assurance, SQA),一种用于管理非确定性智能体基础设施的控制平面原语。SQA将提案表示为绑定密码学证据链的声明式执行契约,并路由至一组多样化的只读沙箱验证代理。SQA通过风险自适应的多数决谓词聚合其判断,该谓词强制模型与原型多样性,基于校准的保证分数调整权重,并尊重原型特定否决权。获准提案仅可通过主权执行门执行。我们在云原生控制平面中实现SQA,形式化了非确定性验证者的相关认知失效模型。在500个受基础设施启发的变更场景中,排除模糊场景后,持有保留集上的安全结果表明,单代理验证的不安全批准率从18.5%降至0.3%,跨研究风险桶的平均验证延迟增加1.45–4.12秒。
原文摘要 · Abstract (English)
As large language model (LLM) agents are integrated into autonomous cloud operations, distributed systems face a semantic reliability problem: proposer agents can generate production mutations, such as modifying IAM policies, opening firewall security groups, or executing data exports, that are syntactically valid and statically authorized but operationally unsafe. Classical distributed consensus protocols replicate deterministic state transitions but do not evaluate the safety of the proposed intent. To address this gap, we introduce Semantic Quorum Assurance (SQA), a control-plane primitive for governing non-deterministic agentic infrastructure. SQA represents proposals as declarative execution contracts bound to cryptographic evidence chains and routes them to a diverse panel of read-only, sandboxed validator agents. SQA aggregates their judgments under a risk-adaptive quorum predicate that enforces model and archetype diversity, adjusts weights based on calibrated assurance scores, and respects archetype-specific vetoes. Admitted proposals execute only through a sovereign execution gate. We instantiate SQA in a cloud-native control plane and formalize a correlated cognitive failure model for non-deterministic validators. On 500 infrastructure-inspired mutation scenarios, with safety results reported on held-out safe/unsafe trials excluding ambiguous scenarios, SQA reduces unsafe approval from 18.5% for single-agent validation to 0.3% while adding median validation latency of 1.45--4.12 seconds across the studied risk buckets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。