为人工智能建立正式的测量理论,让评估更可比、可量化、可监管。
Towards Measurement Theory for Artificial Intelligence
- 构建分层测量框架,区分直接与间接可观测指标。
- 将前沿AI评估与工程安全中的风险分析方法连接起来。
- 揭示能力定义依赖于测量方式,适合研究者与监管者阅读。
我们提出并概述了一个关于人工智能测量的正式理论计划。我们认为,形式化人工智能的测量将使研究人员、从业者和监管者能够:(i) 比较不同系统及其评估方法;(ii) 将前沿AI评估与工程与安全科学中的定量风险分析技术相联系;(iii) 突显所谓人工智能能力的界定依赖于所采用的测量操作与量纲。我们勾勒了一个分层测量架构,区分直接与间接可观测变量,并指出这些要素如何为统一、可校准的人工智能现象分类体系提供路径。
原文摘要 · Abstract (English)
We motivate and outline a programme for a formal theory of measurement of artificial intelligence. We argue that formalising measurement for AI will allow researchers, practitioners, and regulators to: (i) make comparisons between systems and the evaluation methods applied to them; (ii) connect frontier AI evaluations with established quantitative risk analysis techniques drawn from engineering and safety science; and (iii) foreground how what counts as AI capability is contingent upon the measurement operations and scales we elect to use. We sketch a layered measurement stack, distinguish direct from indirect observables, and signpost how these ingredients provide a pathway toward a unified, calibratable taxonomy of AI phenomena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。