arXiv:2604.21854cs.AI2026-04被引 1

为高风险AI系统提供可量化的安全认证方法,解决监管无标准、验证无工具的难题。

Bounding the Black Box: A Statistical Certification Framework for AI Risk Regulation

  • 基于航空认证模式,分两阶段设定风险阈值并统计验证失败率上限
  • 无需模型内部信息,可对任意架构的黑箱系统进行可审计的风险上界计算
  • 适用于需合规的金融、自动驾驶等高风险场景,助力开发者承担主体责任

人工智能已用于决定贷款发放、刑事调查标记及自动驾驶紧急制动。各国政府相继出台监管框架:欧盟《人工智能法案》、NIST风险管理框架、欧洲委员会公约均要求高风险系统在部署前证明安全性。然而,这些法规缺乏对“可接受风险”的量化定义,也未提供技术手段验证系统是否达标。监管架构已建立,但验证工具缺失。这一差距并非理论问题——随着欧盟《人工智能法案》进入全面执法阶段,开发者面临强制合规评估,却无成熟方法生成量化安全证据,而最需监管的黑箱统计推断系统又难以进行白盒分析。本文填补此空白,提出基于航空认证范式的两阶段框架:第一阶段,权威机构明确定义可接受失败概率δ与操作输入域ε;第二阶段,通过RoMA和gRoMA统计验证工具,无需访问模型内部,即可计算系统真实失败率的确定性上界,支持任意模型架构。该证书满足现有监管要求,将责任前置至开发者,并可融入现有法律体系。

原文摘要 · Abstract (English)

Artificial intelligence now decides who receives a loan, who is flagged for criminal investigation, and whether an autonomous vehicle brakes in time. Governments have responded: the EU AI Act, the NIST Risk Management Framework, and the Council of Europe Convention all demand that high-risk systems demonstrate safety before deployment. Yet beneath this regulatory consensus lies a critical vacuum: none specifies what ``acceptable risk'' means in quantitative terms, and none provides a technical method for verifying that a deployed system actually meets such a threshold. The regulatory architecture is in place; the verification instrument is not. This gap is not theoretical. As the EU AI Act moves into full enforcement, developers face mandatory conformity assessments without established methodologies for producing quantitative safety evidence - and the systems most in need of oversight are opaque statistical inference engines that resist white-box scrutiny. This paper provides the missing instrument. Drawing on the aviation certification paradigm, we propose a two-stage framework that transforms AI risk regulation into engineering practice. In Stage One, a competent authority formally fixes an acceptable failure probability $δ$ and an operational input domain $\varepsilon$ - a normative act with direct civil liability implications. In Stage Two, the RoMA and gRoMA statistical verification tools compute a definitive, auditable upper bound on the system's true failure rate, requiring no access to model internals and scaling to arbitrary architectures. We demonstrate how this certificate satisfies existing regulatory obligations, shifts accountability upstream to developers, and integrates with the legal frameworks that exist today.

AI监管风险认证黑箱验证合规框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。