arXiv:2502.03253cs.CL2025-02被引 3

比较人类与大模型在创意评价中的思维差异,发现两者关注点不同。

How do Humans and Language Models Reason About Creativity? A Comparative Analysis

  • 通过解释评分理由,分析人类与模型对创意的判断逻辑。
  • 人类更关注方案稀有性,而模型侧重想法的陌生度和语义距离。
  • 模型受示例影响更大,评价趋于同质化,适合评估系统设计参考。

科学与工程领域的创意评估正越来越多地结合人工与人工智能判断,但其背后的认知过程与偏见仍不清晰。我们开展了两项实验,研究提供示例解决方案及评分对创意评价的影响。在第一项研究中,72名具备科学或工程训练背景的专家被分为两组:一组收到示例(example),另一组未收到(no example)。细粒度标注显示,无示例专家更多使用比较性语言(如“更好/更差”),并更强调方案的稀有性,暗示其依赖记忆检索进行对比。第二项研究对前沿大语言模型进行平行分析发现,模型在评分时优先考虑想法的稀有性和陌生度,表明其评价基于语义相似性。在示例条件下,尽管模型预测真实原始分的准确率提升,但陌生度、稀有性和巧妙性与原始性之间的相关性显著上升,最高达0.99,表明模型对各维度的评价趋于同质化。研究揭示了人类与AI在创意推理上的根本差异,并提示不同群体对创意评判的优先级存在分歧。

原文摘要 · Abstract (English)

Creativity assessment in science and engineering is increasingly based on both human and AI judgment, but the cognitive processes and biases behind these evaluations remain poorly understood. We conducted two experiments examining how including example solutions with ratings impact creativity evaluation, using a finegrained annotation protocol where raters were tasked with explaining their originality scores and rating for the facets of remoteness (whether the response is "far" from everyday ideas), uncommonness (whether the response is rare), and cleverness. In Study 1, we analyzed creativity ratings from 72 experts with formal science or engineering training, comparing those who received example solutions with ratings (example) to those who did not (no example). Computational text analysis revealed that, compared to experts with examples, no-example experts used more comparative language (e.g., "better/worse") and emphasized solution uncommonness, suggesting they may have relied more on memory retrieval for comparisons. In Study 2, parallel analyses with state-of-the-art LLMs revealed that models prioritized uncommonness and remoteness of ideas when rating originality, suggesting an evaluative process rooted around the semantic similarity of ideas. In the example condition, while LLM accuracy in predicting the true originality scores improved, the correlations of remoteness, uncommonness, and cleverness with originality also increased substantially -- to upwards of $0.99$ -- suggesting a homogenization in the LLMs evaluation of the individual facets. These findings highlight important implications for how humans and AI reason about creativity and suggest diverging preferences for what different populations prioritize when rating.

创意评估大模型人类对比认知差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。