arXiv:2504.15784cs.CLcs.AI2025-04EMNLP被引 26

用参考文本评估大模型写作创意,比人工更准更省

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

  • 以优质参考文本为基准,按量表评分生成文本创意
  • 与人类评估一致率达75%,提升15个百分点
  • 适合评测写作类大模型创意能力

创意写作是大语言模型的重要能力,广泛应用于文学、叙事及各类创造性领域。然而,机器生成文本的创意评估仍面临挑战,现有方法或依赖昂贵的人工标注,或难以贴近人类判断。本文提出一种基于托兰斯创意写作测试(TTCW)的自动化评估方法,将创意视为可评价的产品。该方法采用参考文本驱动的李克特量表式评分,从多个维度对生成文本与高质量参考文本进行对比打分。实验表明,该方法显著提升了大模型评估结果与人类评判的一致性,配对准确率达到0.75,较此前方法提升15%。

原文摘要 · Abstract (English)

Creative writing is a key capability of Large Language Models (LLMs), with potential applications in literature, storytelling, and various creative domains. However, evaluating the creativity of machine-generated texts remains a significant challenge, as existing methods either rely on costly manual annotations or fail to align closely with human assessments. In this paper, we propose an effective automated evaluation method based on the Torrance Test of Creative Writing (TTCW), which evaluates creativity as product. Our method employs a reference-based Likert-style approach, scoring generated creative texts relative to high-quality reference texts across various tests. Experimental results demonstrate that our method significantly improves the alignment between LLM evaluations and human assessments, achieving a pairwise accuracy of 0.75 (+15\%).

创意评估大模型评测自动评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。