arXiv:2608.03866cs.AIcs.SY2026-08

为工业大模型建议的合理性设计安全评估框架

ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories

  • 基于版本化工厂档案,三重检查建议是否合规
  • 明确拒绝无证据支持或越权的决策建议
  • 适合研究者评估工业AI建议安全性

本白皮书提出ADMITBench,一个用于评估工业大模型建议可行性的参考框架。该框架通过版本化的安全治理评估契约,检查某项建议是否具备充分证据支持、符合既定权限与程序要求,并满足所选评估配置文件中编码的工厂特异性后果约束。其中,'安全治理'指通过显式、非补偿性检查确定资格,不意味着评估者、模型或工厂已通过安全认证。Release 0.1.0为技术与研究评估提供的公开参考实现,不代表可执行于实际生产环境。

原文摘要 · Abstract (English)

This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework implements a versioned, safety-governed evaluation contract that checks whether a recommendation is supported by the available evidence, permitted under the stated authority and procedure, and acceptable under the plant-specific consequence checks encoded in the selected evaluation profile. In this report, \emph{safety-governed} means that eligibility is determined through explicit, non-compensatory checks derived from a versioned plant profile; it does not mean that the evaluator, model, or plant has been safety-certified. Release 0.1.0 is a public reference implementation for technical and research evaluation, not an authorisation for physical execution.

大模型评估工业AI安全治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。