SAGE构建安全优先的生成式AI全生命周期防护体系,防止重大滥用风险。
SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

- 采用签名发布清单与多维度检测,实现安全优先的授权分离架构
- 实测840次调用中仅449次成功判断,有害合规率较低且受协议约束
- 适合关注AI安全治理、系统级防护设计的研究者与开发者
高影响力生成式AI的灾难性滥用是全生命周期管控问题,而非仅靠提示过滤可解决。SAGE是一种以安全为先、授权分离的架构,将可信的灾难性启用风险作为准入前提,优先于效能、延迟或商业目标考量。该架构融合签名发布清单、多样检测器、鲁棒风险边界、最低风险默认设置、输出检查、三值监控、受保护审计链、隔离机制与回滚功能。形式化证明确立了安全优先性、保守检测边界、单调释放门控、防篡改记录及授权切割;两个PRISM抽象在显式假设下验证了授权分离与生命周期不变量。一项冻结的、厂商对称的研究向四个GPT、四个Claude及两个Gemini快照各发送84例,共840次调用产生794个目标响应、46次提供方错误,449次成功判断覆盖375个响应,八个快照实现完整领域覆盖。有害合规估计值偏低,差异主要源于良性效用与安全引导。七组多重性校正后的对比(涉及Claude、Gemini或GPT-5快照与GPT-5 mini/nano)成立,但未发现Claude或Gemini快照与GPT-5或GPT-5.5之间的有效对比。观测到的有害合规范围是在每提示仅一次生成、无工具、无检索、无历史、无人工裁决的协议约束下保守估计,非实际辅助能力上限。预注册扩展说明如何通过锁定划分、重复采样、多轮对话与沙箱工具条件及领域专家评分测试更广的最佳-最差差距。
原文摘要 · Abstract (English)
High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authorization-separated architecture in which credible catastrophic-enablement risk constrains admissibility before utility, latency, or commercial objectives are considered. It combines signed release manifests, diverse detectors, robust risk envelopes, least-risk defaults, output checking, three-valued monitoring, protected audit chains, containment, and rollback. Formal results establish safety priority, conservative detector bounds, monotone release gating, tamper-evident records, and an authorization cut; two PRISM abstractions verify authorization separation and lifecycle invariants under explicit assumptions. A frozen, vendor-symmetric study sent 84 cases to each of four GPT, four Claude, and two Gemini snapshots: 840 calls yielded 794 target responses, 46 provider errors, and 449 successful judgments covering 375 responses. Eight snapshots had complete judged domain coverage. Harmful-compliance estimates were low; variation arose mainly from benign utility and safe redirection. Seven multiplicity-adjusted contrasts involving Claude, Gemini, or GPT-5 snapshots and the GPT-5 mini and GPT-5 nano snapshots were supported, while no tested contrast between the Claude or Gemini snapshots and GPT-5 or GPT-5.5 survived correction. The observed harmful-compliance range is a conservative, protocol-bound view from one generation per prompt with no tools, retrieval, history, or human adjudication; it is not an upper bound on operational assistance. A preregistered extension specifies how to test a wider best-worst gap using a locked split, repeated sampling, multi-turn and sandboxed-tool conditions, and domain-expert scoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。