arXiv:2503.06378cs.AIcs.CL2025-03被引 37

提出可解释且能预测AI表现的新评估体系,突破传统基准局限。

General Scales Unlock AI Evaluation with Explanatory and Predictive Power

  • 构建18个不饱和的通用量表与评分规则,量化任务需求与模型能力
  • 在63个任务上验证,对大模型性能预测准确率显著优于黑箱基线
  • 揭示模型规模、思维链等对知识、元认知的影响,适合评估部署前风险

确保AI安全高效应用需理解并预测其在新任务上的表现,涵盖从科学挑战到职场变革。现有基准虽推动进展,但对通用AI系统缺乏解释力与预测力,因任务间迁移性差。本文提出通用量表评估体系,可揭示基准真实测量内容,提取模型能力画像,并预测其在分布内/外新任务的表现。基于18项新设计的评分规则,该全自动方法对15个大语言模型和63个任务进行分析,通过对比任务需求与模型能力,揭示不同基准的敏感性与特异性,以及模型规模、思维链与蒸馏对知识、元认知和推理的影响。令人惊讶的是,基于量表的需求水平可在实例级实现高预测力,尤其在分布外场景(新任务、新基准)中显著优于基于嵌入或微调的黑箱基线。所提出的量表、规则、测试集、技术与结果为AI评估带来重大进展,支撑未来AI的可靠部署。(协作平台:https://kinds-of-intelligence-cfi.github.io/ADELE/)

原文摘要 · Abstract (English)

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activities. So far, benchmarking has guided progress in AI, but it has offered limited explanatory and predictive power for general-purpose AI systems, given the low transferability across diverse tasks. In this paper, we introduce general scales for AI evaluation that can explain what common AI benchmarks really measure, extract ability profiles of AI systems, and predict their performance for new task instances, in- and out-of-distribution. Our fully-automated methodology builds on 18 newly-crafted rubrics that place instance demands on general scales that do not saturate. Illustrated for 15 large language models and 63 tasks, high explanatory power is unleashed from inspecting the demand and ability profiles, bringing insights on the sensitivity and specificity exhibited by different benchmarks, and how knowledge, metacognition and reasoning are affected by model size, chain-of-thought and distillation. Surprisingly, high predictive power at the instance level becomes possible using these demand levels, providing superior estimates over black-box baseline predictors based on embeddings or finetuning, especially in out-of-distribution settings (new tasks and new benchmarks). The scales, rubrics, battery, techniques and results presented here represent a major step for AI evaluation, underpinning the reliable deployment of AI in the years ahead. (Collaborative platform: https://kinds-of-intelligence-cfi.github.io/ADELE.)

AI评估可解释性预测能力大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。