arXiv:2505.19236cs.CL2025-05中稿 · ICLR被引 11

构建100万+创意文本数据集,用AI自动评估文本创造力。

Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator

  • 基于共享上下文指令的配对评估框架,提升评价一致性。
  • 在100万+合成数据上训练的CrEval模型,与人类判断高度一致。
  • 适合需要提升AI创造力的研究者和开发者使用。

创意评估仍是大语言模型(LLM)面临的挑战性课题。现有评估严重依赖低效且昂贵的人工评判,阻碍了机器创造力的提升。尽管已有自动化方法,包括心理测试、启发式或提示驱动的方案,但普遍存在泛化能力差或与人类判断不一致的问题。为此,我们提出一种新颖的配对比较框架,通过共享上下文指令提升评估一致性。我们构建了CreataSet,一个包含100万+人工级和100万+合成创意指令-响应对的大规模数据集,覆盖多样化开放域任务。基于该数据集训练的LLM评估器CrEval,在与人类判断的一致性上显著优于现有方法。实验表明,融合人工与合成数据对训练鲁棒评估器至关重要,并验证了CrEval在提升LLM创造力方面的实际应用价值。

原文摘要 · Abstract (English)

Creativity evaluation remains a challenging frontier for large language models (LLMs). Current evaluations heavily rely on inefficient and costly human judgments, hindering progress in enhancing machine creativity. While automated methods exist, ranging from psychological testing to heuristic- or prompting-based approaches, they often lack generalizability or alignment with human judgment. To address these issues, we propose a novel pairwise-comparison framework for assessing textual creativity that leverages shared contextual instructions to improve evaluation consistency. We introduce CreataSet, a large-scale dataset with 100K+ human-level and 1M+ synthetic creative instruction-response pairs spanning diverse open-domain tasks. Through training on CreataSet, we develop an LLM-based evaluator named CrEval. CrEval demonstrates remarkable superiority over existing methods in alignment with human judgments. Experimental results underscore the indispensable significance of integrating both human and synthetic data to train highly robust evaluators, and showcase the practical utility of CrEval in boosting the creativity of LLMs.

创意评估大模型评测数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。