arXiv:2410.04265cs.CL2024-10被引 47

用网页文本重构法量化语言模型的创造性,发现人类远超AI。

AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text

  • 通过动态规划算法从网络文本中匹配原文片段,评估文本创造性。
  • 人类作者平均创造性比LLM高66.2%,对齐使模型创造性下降30.1%。
  • 该指标可有效识别机器生成文本,零样本检测性能超越现有系统。

创造力长期被视为人工智能难以模仿的人类智能核心特征。然而,以ChatGPT为代表的大型语言模型(LLMs)的兴起,引发了关于AI是否能媲美甚至超越人类创造力的讨论。本文提出CREATIVITY INDEX,作为首个通过从网络文本片段重构来量化文本语言创造性的方法。该指标基于假设:大模型看似惊人的创造性,很大程度上源于网络上的原创人类文本。为高效计算,我们引入DJ SEARCH——一种新颖的动态规划算法,可搜索文档片段与网络文本的完全或近似匹配。实验表明,专业人类作者的CREATIVITY INDEX平均比LLM高出66.2%,对齐操作使模型创造性平均降低30.1%。此外,海明威等杰出作家的创造性显著高于其他人类写作者。最后,我们证明该指数可作为零样本机器文本检测的有效标准,在六项任务中五项优于最强监督系统GhostBuster,且在零样本场景下超越最强现有系统DetectGPT达30.2%。

原文摘要 · Abstract (English)

Creativity has long been considered one of the most difficult aspect of human intelligence for AI to mimic. However, the rise of Large Language Models (LLMs), like ChatGPT, has raised questions about whether AI can match or even surpass human creativity. We present CREATIVITY INDEX as the first step to quantify the linguistic creativity of a text by reconstructing it from existing text snippets on the web. CREATIVITY INDEX is motivated by the hypothesis that the seemingly remarkable creativity of LLMs may be attributable in large part to the creativity of human-written texts on the web. To compute CREATIVITY INDEX efficiently, we introduce DJ SEARCH, a novel dynamic programming algorithm that can search verbatim and near-verbatim matches of text snippets from a given document against the web. Experiments reveal that the CREATIVITY INDEX of professional human authors is on average 66.2% higher than that of LLMs, and that alignment reduces the CREATIVITY INDEX of LLMs by an average of 30.1%. In addition, we find that distinguished authors like Hemingway exhibit measurably higher CREATIVITY INDEX compared to other human writers. Finally, we demonstrate that CREATIVITY INDEX can be used as a surprisingly effective criterion for zero-shot machine text detection, surpassing the strongest existing zero-shot system, DetectGPT, by a significant margin of 30.2%, and even outperforming the strongest supervised system, GhostBuster, in five out of six domains.

语言模型创造性评估文本检测定量分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。