arXiv:2606.11762cs.CLcs.AI2026-06ACL被引 1

提出可跨任务评估大模型创造力的自动化框架,摆脱领域依赖。

Automated Creativity Evaluation of Language Models Across Open-Ended Tasks

论文配图:Automated Creativity Evaluation of Language Models Across Open-Ended Tasks
图 1 · 摘自论文原文
  • 用语义熵衡量创意新颖性与多样性,无需参考答案。
  • 多智能体检索裁判系统提升任务完成度评估效率超60%。
  • 在三类开放任务中验证,适配不同规模与推理能力的模型。

大语言模型在理解、推理和生成方面取得显著进展,激发了对其创造潜力的关注。实现这一潜力需系统化、可扩展的创造力评估方法。然而,现有多数创造力指标与特定任务紧密耦合,嵌入领域假设,限制了可扩展性和通用性。为此,我们提出一种自动化、领域无关的框架,用于量化大模型在开放任务中的创造力。该方法将评估机制与具体创作任务解耦,实现可扩展的跨任务评估。发散创造力通过语义熵衡量,该指标无参考、鲁棒性强,经人类标注、基于大模型的新颖性判断及基线多样性度量验证。收敛创造力采用新型基于检索的多智能体裁判框架,实现上下文敏感的任务完成度评估,效率提升超过60%。我们在三个质异领域(问题解决:MacGyver;研究构思:HypoGen;创意写作:BookMIA)中,使用广泛的LLMs进行验证。实证结果表明,该框架能可靠捕捉创造力的关键维度——新颖性、多样性与任务契合度,并揭示模型规模、温度、更新时间与推理能力对创造性表现的影响。本工作建立了一个可复现、通用的自动化创造力评估标准,为可扩展基准测试铺路,加速创意AI发展。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable progress in language understanding, reasoning, and generation, sparking growing interest in their creative potential. Realizing this potential requires systematic and scalable methods for evaluating creativity across diverse tasks. However, most existing creativity metrics are tightly coupled to specific tasks, embedding domain assumptions into the evaluation process, and limiting scalability and generality. To address this gap, we introduce an automated, domain-agnostic framework for quantifying LLM creativity across open-ended tasks. Our approach separates the measurement apparatus from the creative task itself, enabling scalable, task-agnostic assessment. Divergent creativity is measured using semantic entropy, a reference-free and robust metric for novelty and diversity, validated against human annotations, LLM-based novelty judgments and baseline diversity measures. Convergent creativity is assessed via a novel retrieval-based multi-agent judge framework that delivers context-sensitive evaluation of task fulfilment with over 60% improved efficiency. We validate our framework in three qualitatively distinct domains: problem-solving (MacGyver), research ideation (HypoGen), and creative writing (BookMIA), using a broad suite of LLMs. Empirical results show that our framework reliably captures key facets of creativity, including novelty, diversity, and task fulfilment, and reveal how model properties, such as size, temperature, recency, and reasoning, impact creative performance. Our work establishes a reproducible and generalizable standard for automated LLM creativity evaluation, paving the way for scalable benchmarking and accelerating progress in creative AI.

创造力评估大模型自动化评测开放任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。