arXiv:2605.10601cs.AI2026-05

AI部署应靠可验证的监管机制,而非追求模型内部解释。

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime

  • 用可独立验证的授权机制替代对模型内部的过度解释要求。
  • 53个百分点差距显示理解不等于可控,仅9.0%的医疗AI文档含上市后监测。
  • 适合政策制定者、监管机构和高风险领域AI开发者参考。

在医疗、信贷、就业和刑事司法等敏感领域,AI部署常因无法解释模型内部机制而被视为不安全。这种做法导致过度依赖机制可解释性来解决本不属于其范畴的部署问题。我们主张采用校准的验证机制:授权应针对具体应用场景,具备独立可验证性、上线后监控、责任追究、申诉与撤销能力。原因有二:其一,模型能力在相近任务间差异显著,授权必须绑定具体用途;其二,社会长期通过资质认证、持续监控、责任追溯、申诉与撤销等方式管理不透明的专业知识。近期证据表明,内部表征与输出修正之间存在53个百分点的差距,说明理解未必能转化为有效控制;一项范围综述发现,仅9.0%的FDA批准的AI/ML医疗设备文档包含前瞻性上市后监测研究。为此,我们提出‘验证覆盖率’(Verification Coverage),一个包含六个组件的可报告标准,并设置最低组成规则,作为模型卡、排行榜与监管披露中与能力评分并列的核心指标。

原文摘要 · Abstract (English)

AI deployment in sensitive domains such as health care, credit, employment, and criminal justice is often treated as unsafe to authorize until model internals can be explained. This often leads to an excessive reliance on mechanistic interpretability to address a deployment challenge beyond its intended scope. We argue that the gate should instead be calibrated verification: authorization should be domain-scoped, independently checkable, monitored after release, accountable, contestable, and revocable. The reason is twofold. First, model capability is uneven across nearby tasks, so authorization must attach to a specific use rather than to a model in general. Second, societies have long governed opaque expertise through credentials, monitoring, liability, appeal, and revocation rather than mechanism-level explanation. Recent evidence reinforces this distinction between mechanistic understanding and deployment authority: a 53-percentage-point gap between internal representations and output correction shows that understanding may not translate into action, while one scoping review found that only 9.0% of FDA-approved AI/ML device documents contained a prospective post-market surveillance study. We propose Verification Coverage, a six-component reportable standard with a minimum-composition rule, as the metric that should sit beside capability scores in model cards, leaderboards, and regulatory disclosures.

AI治理部署安全验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。