arXiv:2411.15560cs.AIcs.CL2024-11被引 7

大模型对创意测试答案的评价高度一致,可信赖用于自动创意评估。

Do LLMs Agree on the Creativity Evaluation of Alternative Uses?

  • 用四种顶级大模型评估不同来源的创意回答,检验其一致性。
  • 跨模型相关性超0.7,与权威标准相关性达0.77以上。
  • 模型不偏爱自己生成的答案,具备公正性,适合自动化创意评测。

本文研究大语言模型(LLMs)在评估替代用途测试(AUT)中创造力时是否达成共识。尽管大模型被广泛用于创意内容评估,但以往研究多聚焦单一模型对同类生成内容的评价。本文通过一个由创造性等级(普通、有创意、高度创意)标注的基准数据集,测试四种先进大模型对自身及其他模型生成的回答进行评分和排序的能力。实验采用全面与分段两种评估设置。结果显示,各模型间具有高一致性,平均斯皮尔曼相关系数超过0.7,与权威标准的相关性达0.77以上,表明大模型在替代用途创造力评估上具有高度可靠性。值得注意的是,模型并未偏好自身输出,对其他模型生成的内容给出相似评分或排名。这些发现表明大模型在创造力评估中表现出公正性与高度一致性,为自动化创意评估提供了有力支持。

原文摘要 · Abstract (English)

This paper investigates whether large language models (LLMs) show agreement in assessing creativity in responses to the Alternative Uses Test (AUT). While LLMs are increasingly used to evaluate creative content, previous studies have primarily focused on a single model assessing responses generated by the same model or humans. This paper explores whether LLMs can impartially and accurately evaluate creativity in outputs generated by both themselves and other models. Using an oracle benchmark set of AUT responses, categorized by creativity level (common, creative, and highly creative), we experiment with four state-of-the-art LLMs evaluating these outputs. We test both scoring and ranking methods and employ two evaluation settings (comprehensive and segmented) to examine if LLMs agree on the creativity evaluation of alternative uses. Results reveal high inter-model agreement, with Spearman correlations averaging above 0.7 across models and reaching over 0.77 with respect to the oracle, indicating a high level of agreement and validating the reliability of LLMs in creativity assessment of alternative uses. Notably, models do not favour their own responses, instead they provide similar creativity assessment scores or rankings for alternative uses generated by other models. These findings suggest that LLMs exhibit impartiality and high alignment in creativity evaluation, offering promising implications for their use in automated creativity assessment.

大模型创意评估一致性自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。