用社区结构分析发现,大模型生成文本仍与人类语言有本质差异。
Does a Large Language Model Really Speak in Human-Like Language?
- 通过对比人类与大模型生成文本的潜在社区结构差异
- 发现大模型生成文本始终与人类写作存在显著区别
- 适合关注大模型语言真实性的研究人员阅读
大型语言模型(LLMs)近期兴起,因其生成高度自然、类人化文本的能力而备受关注。本研究在假设检验框架下,比较了大模型生成文本与人类写作文本之间的潜在社区结构。具体分析三组文本:原始人类写作文本($/mathcal{O}$)、其大模型改写版本($/mathcal{G}$)以及由$/mathcal{G}$再次改写得到的双改写集($/mathcal{S}$)。研究聚焦两个核心问题:(1) $/mathcal{O}$与$/mathcal{G}$之间的社区结构差异是否与$/mathcal{G}$与$/mathcal{S}$之间的差异相当?(2) 当调整控制文本多样性的大模型参数时,$/mathcal{G}$是否更接近$/mathcal{O}$?基于若大模型生成文本真正类人,则两对文本(原-改写)间的差距应相似的假设,提出一种统计假设检验框架。该框架利用各数据集间因改写关系而存在的对应部分,实现跨数据集位置映射,使两组文本可基于第三组文本的空间进行量化比较。结果表明,GPT生成文本与人类文本在社区结构上仍存在明显差异。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently emerged, attracting considerable attention due to their ability to generate highly natural, human-like text. This study compares the latent community structures of LLM-generated text and human-written text within a hypothesis testing procedure. Specifically, we analyze three text sets: original human-written texts ($\mathcal{O}$), their LLM-paraphrased versions ($\mathcal{G}$), and a twice-paraphrased set ($\mathcal{S}$) derived from $\mathcal{G}$. Our analysis addresses two key questions: (1) Is the difference in latent community structures between $\mathcal{O}$ and $\mathcal{G}$ the same as that between $\mathcal{G}$ and $\mathcal{S}$? (2) Does $\mathcal{G}$ become more similar to $\mathcal{O}$ as the LLM parameter controlling text variability is adjusted? The first question is based on the assumption that if LLM-generated text truly resembles human language, then the gap between the pair ($\mathcal{O}$, $\mathcal{G}$) should be similar to that between the pair ($\mathcal{G}$, $\mathcal{S}$), as both pairs consist of an original text and its paraphrase. The second question examines whether the degree of similarity between LLM-generated and human text varies with changes in the breadth of text generation. To address these questions, we propose a statistical hypothesis testing framework that leverages the fact that each text has corresponding parts across all datasets due to their paraphrasing relationship. This relationship enables the mapping of one dataset's relative position to another, allowing two datasets to be mapped to a third dataset. As a result, both mapped datasets can be quantified with respect to the space characterized by the third dataset, facilitating a direct comparison between them. Our results indicate that GPT-generated text remains distinct from human-authored text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。