arXiv:2412.00530cs.AIcs.CL2024-12被引 3

用网络模型分析文本创意,发现GPT-3.5评分与人类差异显著。

Forma mentis networks predict creativity ratings of short texts via interpretable artificial intelligence in human and GPT-simulated raters

  • 构建文本思维网络,提取语义、句法与情感特征
  • GPT-3.5更偏好自己生成的故事,评分模式不同
  • 解释性AI揭示其评分机制缺乏人类创意复杂性

创造力是人类认知的核心能力。我们使用文本思维网络(TFMN)从约一千篇人类与GPT-3.5生成的故事中提取网络(语义/句法关联)和情感特征。通过可解释人工智能(XAI),检验这些特征是否能解释人类及GPT-3.5对创意的评分。采用XGBoost分析三种场景:(i) 人类评人类故事,(ii) GPT-3.5评人类故事,(iii) GPT-3.5评自身生成故事。结果表明,GPT-3.5的评分不仅与人类相关性不同,且在特征模式上存在显著差异。它更偏爱自身生成的内容,对人类故事的评价方式不同于人类。基于SHAP的特征重要性分析显示:(i) 网络特征对人类评分及GPT-3.5评人类故事更具预测力;(ii) 在评分自身生成故事时,情感特征比语义/句法结构作用更大。定量结果凸显了GPT-3.5在契合人类创意评估上的关键局限,提醒在用于创意内容评估与生成时需谨慎。

原文摘要 · Abstract (English)

Creativity is a fundamental skill of human cognition. We use textual forma mentis networks (TFMN) to extract network (semantic/syntactic associations) and emotional features from approximately one thousand human- and GPT3.5-generated stories. Using Explainable Artificial Intelligence (XAI), we test whether features relative to Mednick's associative theory of creativity can explain creativity ratings assigned by humans and GPT-3.5. Using XGBoost, we examine three scenarios: (i) human ratings of human stories, (ii) GPT-3.5 ratings of human stories, and (iii) GPT-3.5 ratings of GPT-generated stories. Our findings reveal that GPT-3.5 ratings differ significantly from human ratings not only in terms of correlations but also because of feature patterns identified with XAI methods. GPT-3.5 favours 'its own' stories and rates human stories differently from humans. Feature importance analysis with SHAP scores shows that: (i) network features are more predictive for human creativity ratings but also for GPT-3.5's ratings of human stories; (ii) emotional features played a greater role than semantic/syntactic network structure in GPT-3.5 rating its own stories. These quantitative results underscore key limitations in GPT-3.5's ability to align with human assessments of creativity. We emphasise the need for caution when using GPT-3.5 to assess and generate creative content, as it does not yet capture the nuanced complexity that characterises human creativity.

创意评估GPT-3.5可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。