arXiv:2604.11065cs.AI2026-04被引 2

提出可验证的AI治理新范式,确保决策过程透明可信。

AI Integrity: A New Paradigm for Verifiable AI Governance

  • 构建四层权威栈模型,涵盖价值、认知、信源与数据标准
  • 定义六项核心指标,量化评估推理过程的完整性
  • 适合关注AI可解释性与可信治理的研究者与政策制定者

人工智能系统在医疗、法律、国防和教育等关键领域日益影响重大决策,但现有的治理范式——人工智能伦理、安全与对齐——均存在共同缺陷:仅评估结果而无法验证推理过程。本文提出AI完整性(AI Integrity)概念,指一个AI系统的权威栈(即其分层的价值观、认知标准、信源偏好与数据选择标准)不受腐败、污染、操纵与偏见影响,并可被验证。我们区分了该范式与现有三类治理框架的不同,提出四层权威栈模型:规范层(基于舒瓦茨基本人类价值观)、认知层(基于沃尔顿论证模式及GRADE/CEBM分级体系)、信源层(基于信源可信度理论)与数据层。明确区分合法级联与权威污染,并识别出‘完整性幻觉’为威胁价值一致性的核心可测量风险。进一步提出PRISM框架作为操作方法,定义六项核心度量指标并规划分阶段研究路线。不同于规定正确价值观的规范框架,AI完整性是一种程序性概念:要求从证据到结论的路径始终透明且可审计,无论系统持何种价值立场。

原文摘要 · Abstract (English)

AI systems increasingly shape high-stakes decisions in healthcare, law, defense, and education, yet existing governance paradigms -- AI Ethics, AI Safety, and AI Alignment -- share a common limitation: they evaluate outcomes rather than verifying the reasoning process itself. This paper introduces AI Integrity, a concept defined as a state in which the Authority Stack of an AI system -- its layered hierarchy of values, epistemological standards, source preferences, and data selection criteria -- is protected from corruption, contamination, manipulation, and bias, and maintained in a verifiable manner. We distinguish AI Integrity from the three existing paradigms, define the Authority Stack as a 4-layer cascade model (Normative, Epistemic, Source, and Data Authority) grounded in established academic frameworks -- Schwartz Basic Human Values for normative authority, Walton argumentation schemes with GRADE/CEBM hierarchies for epistemic authority, and Source Credibility Theory for source authority -- characterize the distinction between legitimate cascading and Authority Pollution, and identify Integrity Hallucination as the central measurable threat to value consistency. We further specify the PRISM (Profile-based Reasoning Integrity Stack Measurement) framework as the operational methodology, defining six core metrics and a phased research roadmap. Unlike normative frameworks that prescribe which values are correct, AI Integrity is a procedural concept: it requires that the path from evidence to conclusion be transparent and auditable, regardless of which values a system holds.

AI治理可验证性推理透明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。