首个统一评估AI通用创造力的基准,揭示模型在不同创意领域表现差异。
AGC-Bench: Measuring Artificial General Creativity
- 基于3101篇文献构建跨领域创意评测集,支持多任务评估
- 发现大模型存在单一创造力因子,解释81.5%变异且独立于通用能力
- 推出可开源评分模型与人类对照数据,推动公平评测
创造力研究长期争论创造力是否具领域特异性以及能否与一般智力分离。这一问题如今也适用于大语言模型(LLMs),但缺乏统一的AI创造力评估基准。本文提出AGC-Bench,一个通过系统性回顾3,101篇文献(筛选出497个基准)构建的通用人工智能创造力评测体系,结合代理机制将异构代码库转化为HELM标准格式。首次发布涵盖78个数据集,覆盖头脑风暴、问题解决、STEM、叙事、隐喻语言和幽默等维度。为缓解LLM作为评判者的偏差,引入裁判反应理论进行心理计量校准,并微调Qwen3-30B模型生成开放权重的AGC-Judge,可在未训练过的任务上稳定评分。结果显示,前沿模型位居榜首,开源模型紧随其后;模型在写作类任务中表现更优,而在科学构思中较弱。实验证明:对83个模型进行因子分析,可提取出一个解释81.5%方差的创造力因子'c',虽与通用知识/推理相关,但可分离;主动提示‘发挥创造力’显著提升性能,远超启用推理能力的效果;在人类匹配子集上,顶尖人类仍领先顶尖模型。项目已开源评测基准、公开排行榜、评分模型及人类数据,旨在实现大规模、可扩展的AI创造力测量。
原文摘要 · Abstract (English)
Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both questions now apply to LLMs, but a unified benchmark of AI creativity remains elusive. We introduce AGC-Bench, an artificial general creativity benchmark built from a systematic review of the AI creativity literature (3,101 papers screened, 497 benchmarks identified), paired with an agentic harness that converts idiosyncratic codebases into HELM-standardized benchmarks. The first release covers 78 datasets spanning brainstorming, problem solving, STEM, narrative, figurative language, and humor. To address bias in LLM-as-judge, we apply Judge Response Theory -- a psychometric calibration of judge leniency/severity; we then fine-tune Qwen3-30B on the bias-corrected ratings of three frontier LLMs to produce AGC-Judge, an open-weight model that robustly scores new creativity benchmarks it was not trained on. Results reveal frontier models at the top of the AGC-Bench leaderboard, with open models close behind. LLMs show different creative strengths, ranking higher on some domains (e.g., writing) than others (e.g., scientific ideation). Extensive experiments yield three main findings. First, applying factor analysis across 83 LLMs, we recover a single creativity factor 'c', analogous to the 'g' factor of general intelligence, that explains 81.5% of variance, related to but separable from general knowledge/reasoning. Second, we show that prompting models to "be creative" boosts their performance far more than enabling reasoning, evidence that the benchmark tracks creativity over general ability. Third, on a human-matched subset, we find the top human still leads the top LLM on creativity. We release AGC-Bench with a public leaderboard, AGC-Judge, and human data as open infrastructure for measuring AI creativity at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。