arXiv:2607.22218cs.CLcs.AI2026-07被引 1

大模型评价创意时,为何有时像人,有时不像?

Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

  • 发现大模型评价创意主要依赖新颖性,忽视社会市场等背景信息
  • 不同大模型标准差异大,部分模型更擅长区分人类认为更创意的想法
  • 当判断需考虑现实背景时,大模型与人类评价明显分歧

尽管大语言模型(LLMs)被广泛用作创意评估工具,但其与人类评价的对齐程度仍不一致。本研究通过三项实验及六种主流大模型,揭示了大模型创意评估的标准及其下游影响。研究1显示,大模型普遍依赖较窄的人类创意评估标准,尤其在新颖性维度上与人类高度一致,但在涉及社会、市场与声誉等上下文信息的维度上差异显著;各模型自身标准存在明显差异。研究2(N=1,103个创意点子)表明,大模型评价与人类评价中度相关,标准更广的模型能更好区分人类认为更具创意的想法。研究3(N=1,195)发现,上下文信息显著影响人类评分,但对大模型评分影响甚微。结果说明:大模型与人类在强调内在特质(如新颖性)时可能对齐,但在需要背景信息时则显著偏离。因此,选择大模型作为评估工具是关键决策,不同模型识别出的创意也不同。

原文摘要 · Abstract (English)

Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative.

创意评估大模型对齐人类对比上下文感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。