发现大模型评估创意原创性时偏好自己生成的内容,但细化后偏差消失。
The Effect of Idea Elaboration on the Automatic Assessment of Idea Originality

- 对比人类与大模型对创意想法的原创性评分,分析4813个回答。
- 未控制细节时,大模型更偏爱自己生成的内容,存在自我偏好偏差。
- 当考虑想法细化程度后,这种偏差消失,提示评估需关注细节层次。
自动系统在创造性任务中评估响应原创性日益普及,可缓解人工评估成本高、疲劳和主观性强等问题,但已有初步证据表明存在自我偏好偏差:自动系统更倾向于选择与自身风格相近的结果而非人类创作。本文研究了大语言模型(LLMs)在发散思维任务中对创意响应原创性的评分是否与人类评审一致。分析了来自高创造力和低创造力人类及ChatGPT-4o的4,813个回答,人类评审员为接受过密集培训的大学生,机器评审系统包括两个针对交替用途任务(AUT)微调的专用系统(OCSAI和CLAUS),以及与人类评审指令相同的ChatGPT-4o。结果证实了大模型存在自我偏好偏差,更青睐人工生成的回答;然而,当分析中控制了想法的细化程度(idea elaboration)后,该偏差消失。这一发现对创造力评估的理论与方法具有重要启示,也为未来研究指明方向。
原文摘要 · Abstract (English)
Automatic systems are increasingly used to assess the originality of responses in creative tasks. They offer a potential solution to key limitations of human assessment (cost, fatigue, and subjectivity), but there is preliminary evidence of a self-preference bias. Accordingly, automatic systems tend to prefer outcomes that are more closely related to their style, rather than to the human one. In this paper, we investigated how Large Language Models (LLMs) align with human raters in assessing the originality of responses in a divergent thinking task. We analysed 4,813 responses to the Alternate Uses Task produced by higher and lower creative humans and ChatGPT-4o. Human raters were two university students who underwent intensive training. Machine raters were two specialised systems fine-tuned on AUT responses and corresponding human ratings (OCSAI and CLAUS) and ChatGPT-4o, which was prompted with the same instructions as human raters. Results confirmed the presence of a self-preference bias in LLMs. Automatic systems tended to privilege artificial responses. However, this self-preference bias disappeared when the analyses controlled for the idea elaboration. We discuss theoretical and methodological implications of these findings by highlighting future directions for research on creativity assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。