arXiv:2512.15783cs.AIcs.LG2025-12中稿 · publication in AI …

建立标准化评估框架,用数据监测部署中AI系统的风险。

Towards AI epidemiology: a measurement standardisation framework for prospective risk detection

  • 构建8个交互字段的结构化评估语法,支持跨模型比较。
  • 实验证明评估结果在指定条件下可重复,可靠性达0.05阈值。
  • 为未来构建'AI流行病学'提供方法基础,适合监管机构使用。

本文提出一种测量标准化框架,将专家与AI的交互压缩为可比较的结构化字段,用于部署中AI系统的事前风险检测,无需访问模型内部。该概念论文定义了框架的语义与统计边界,并制定了实证测试协议。所支持的群体层面结论属于分阶段研究计划,非本文直接成果。测量标准化支撑三大主张:其一为可靠性声明——在有限条件下,大语言模型可生成可靠的证据与政策对齐评估;其二为治理声明——对齐分数能为专家提供实时信号,为机构提供跨任务、模型与领域对齐模式的监控依据;其三为结果验证声明——一旦标准化建立,聚合对齐分数可用于研究其与受监管专业场景下游结果的关联。这引出了“AI流行病学”的可能,即基于相关变量而非机制分析的风险检测,灵感来自流行病学推理。对公开专家-AI语料的最小化应用显示,在规定条件下,评估者两次运行的政策与证据对齐分数一致。大规模评估可靠性有待未来验证。框架包含八项交互字段的定义语法,以及基于配对自举推断、成对AUC的DeLong检验(敏感性检查)、预设单侧非劣效边际0.05和Holm-Bonferroni校正的统计协议。

原文摘要 · Abstract (English)

This paper proposes a measurement standardisation framework that compresses expert-AI interactions into structured, comparable fields for prospective risk detection in deployed AI systems, without access to model internals. This concept paper defines the framework's scope, semantically and statistically, and specifies a protocol for its empirical testing. The population-level claims it is designed to support therefore belong to a staged research programme rather than to results claimed here. Measurement standardisation underpins three claims. The first is a reliability claim: under bounded conditions, large language models can produce reliable, standardised assessments of the evidential and policy alignment of expert-AI interactions. The second is a governance claim: alignment scores give experts an immediate signal during deployment and give institutions a basis for monitoring alignment patterns across mission types, models, and domains. The third is an outcome validation claim: once measurement standardisation is established, aggregate alignment scores could be used to study associations with downstream outcomes in regulated professional settings. This introduces the possibility of an "AI epidemiology", a form of risk detection based on correlated variables instead of mechanistic analysis, inspired by epidemiological reasoning. A minimal application of the protocol to a published expert-AI corpus shows that the judge reproduces its policy and evidential alignment scores across two runs under the specified conditions. Judge reliability at scale remains to be validated in future work. The paper sets out a defined grammar of eight interaction fields, together with a statistical protocol based on paired bootstrap inference, DeLong's test for paired AUCs as a sensitivity check, a pre-specified one-sided non-inferiority margin of 0.05, and Holm-Bonferroni correction.

AI风险评估框架对齐检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。