arXiv:2503.13687cs.CL2025-03被引 3

通过特征分析,可高精度区分人类与GPT生成文本。

Feature Extraction and Analysis for GPT-Generated Text

  • 提取文本风格特征,用机器学习分类器识别生成来源。
  • 足够长的文本下,区分准确率极高。
  • 适合内容审核、学术诚信检测场景使用。

随着GPT等先进自然语言模型的兴起,区分人类撰写与GPT生成文本变得愈发困难且关键,尤其在学术领域。长期存在的抄袭问题如今更添信息真实性疑虑——难以判断所述事实是真实还是虚构。本文系统研究了用于区分人类与GPT生成文本的特征提取与分析方法。通过将机器学习分类器应用于提取的特征,评估各特征在检测中的重要性。结果表明,人类与GPT生成文本在写作风格上存在显著差异,这些差异可通过所提特征有效捕捉。当文本长度足够时,两者可被高精度区分。

原文摘要 · Abstract (English)

With the rise of advanced natural language models like GPT, distinguishing between human-written and GPT-generated text has become increasingly challenging and crucial across various domains, including academia. The long-standing issue of plagiarism has grown more pressing, now compounded by concerns about the authenticity of information, as it is not always clear whether the presented facts are genuine or fabricated. In this paper, we present a comprehensive study of feature extraction and analysis for differentiating between human-written and GPT-generated text. By applying machine learning classifiers to these extracted features, we evaluate the significance of each feature in detection. Our results demonstrate that human and GPT-generated texts exhibit distinct writing styles, which can be effectively captured by our features. Given sufficiently long text, the two can be differentiated with high accuracy.

文本检测GPT特征提取风格分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。