arXiv:2606.21015cs.AI2026-06

AI治理缺了证据充足性这一层,导致风险难控、价值难保。

The AI Evaluability Gap: The Missing Layer for Managing Risk and Sustaining Value

  • 提出‘可评估性’概念,强调证据生成与持续更新能力
  • 区分操作认证与投资认证,分别依赖结构与因果证据
  • 构建六维证据标准,解决治理决策缺乏可信依据的问题

组织在部署AI时面临两大治理挑战:管理风险与维持价值。两者均依赖充分的证据,但当前往往缺乏足够依据支持高置信度决策。本文提出‘AI可评估性鸿沟’概念,指组织缺乏充分证据以支持对风险或价值的高置信决策。现有治理多关注系统属性(如安全、公平、可靠性),却忽视支撑决策的证据基础。我们主张,AI治理包含是否允许运行的操作决策,以及是否持续投入资源的投资决策。为此,提出‘可评估性’——系统长期生成、维护和更新足够证据以支持高置信决策的能力。将治理决策形式化为校准后的置信度Conf(D|E),并定义六项可评估证据属性:可观测性、可归因性、可干预性、可验证性、校准性与时间有效性。框架区分操作认证(依赖结构证据)与投资认证(依赖因果证据)。认为证据充分性是缺失的治理层,填补该鸿沟是管理风险与维持价值的前提。

原文摘要 · Abstract (English)

Organizations deploying AI face two fundamental governance challenges: managing AI risk and sustaining AI value. Both depend on evidence whose sufficiency cannot be taken for granted. We call the shared underlying challenge the AI Evaluability Gap: the condition in which organizations lack sufficient evidence to support high-confidence governance decisions regarding either risk or value. We argue that this gap reflects a category error in current practice. Existing governance approaches focus primarily on properties of systems, such as safety, fairness, reliability, compliance, and value, while paying comparatively little attention to the evidentiary foundations required to justify decisions about those properties. We further argue that AI governance encompasses both operational decisions regarding whether a system may operate and investment decisions regarding whether it merits continued organizational resources. To address this problem, we introduce Evaluability, defined as the capability of a system to generate, maintain, and renew evidence sufficient to support high-confidence governance decisions over time. We formalize governance decisions as functions of calibrated confidence Conf(D|E) and identify six properties of evaluable evidence: observability, attributability, intervenability, verifiability, calibration, and temporal validity. The framework distinguishes Operational Certification, which relies primarily on structural evidence to justify deployment decisions, from Investment Certification, which relies primarily on causal evidence to justify continued resource allocation. We argue that evidence sufficiency is a missing layer of AI governance and that closing the AI Evaluability Gap is a prerequisite for both managing risk and sustaining value in AI-enabled organizations.

AI治理风险控制证据基础可评估性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。