提出AI成熟度自动评估框架RAIL,用大模型分工判断AI项目进展。
RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level
- 构建九级统一成熟度量表,基于自然语言描述即可判定。
- 六维度独立评估+专家仲裁,避免大模型过度乐观误判。
- 适合科研管理、投资决策者快速评估AI项目真实进展。
评估人工智能技术成熟度对投资决策、项目管理和政策监测至关重要,但现有成熟度框架异构且难以自动化:传统技术成熟度等级缺乏AI专属标准,机器学习成熟度等级依赖内部过程数据,而数据/智能成熟度模型采用不可直接比较的量表。本文提出两项贡献:一是将三类框架统一为九级的统一AI成熟度等级(AIRL),基于环境证据阶梯,包含规范性、数据存在性、数据质量、数据合法性、专家知识和算法成熟度六个维度上限,并引入通用性锚定规则与明确赋值准则,使仅凭自然语言描述即可判定成熟度;二是提出RAIL(基于独立LLM专家的成熟度评估),由一个证据代理和六个独立维度代理组成,每个代理均为具备特定任务的大语言模型,其判断经确定性最低规则聚合,再由首席专家在不对称权威下审查,可确认或下调建议,但不可提升至维度上限以上。该方法在多个研究工作分析中验证了结果一致性,避免了单一大模型分类器的高估问题。
原文摘要 · Abstract (English)
Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monitoring, yet the available readiness frameworks are heterogeneous and difficult to apply automatically: the adaptation of Technology Readiness Levels to AI lacks AI-specific gating criteria, the Machine Learning Technology Readiness Levels presuppose access to internal process artifacts, and AI/data readiness dimension models employ scales that resist direct comparison. This paper makes two contributions. First, we unify these three frameworks into the Unified AI Readiness Level (AIRL), a nine-level ordinal scale built on an environmental evidence ladder and complemented by dimensional caps (covering specification, data existence, data quality, data legality, expert knowledge, and algorithmic maturity) together with a generality-anchoring rule and explicit assignment disciplines, so that a readiness level becomes decidable from a natural-language description of the work alone. Second, we propose RAIL (Readiness Assessment via Independent LLM-experts), a panel-of-experts classifier that operationalizes the scale: one evidence agent and six independent dimension agents, each a large language model with a narrowly scoped mandate, deliver verdicts that a deterministic minimum rule aggregates and a chief expert reviews under asymmetric authority, confirming or lowering the panel's recommendation but never raising it above the caps. The method was tested in the analysis of several research works showing consistency and avoiding overestimation from monolithic LLM classifiers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。