用可解释模型识别AI生成的创意写作,准确率达98%。
Interpretable Text Classification Applied to the Detection of LLM-generated Creative Writing
- 采用线性可解释分类器,仅凭单字特征即可区分。
- 准确率93%-98%,关键在于模型能捕捉同义词多样性差异。
- 适合研究生成文本检测或内容可信度验证者阅读。
本文研究如何区分人类撰写的创意小说(小说节选)与由大型语言模型生成的相似文本。结果显示,人类观察者在此二分类任务中表现不佳(接近随机水平),而多种机器学习模型在未见过的测试集上准确率达到0.93至0.98,甚至仅使用短文本样本和单字(一元语法)特征即可实现。因此,我们采用一个内在可解释的线性分类器(测试准确率达0.98),以揭示高准确率背后的原因。分析发现,指示模型生成文本的关键一元语法特征之一是:模型倾向于使用更多样化的同义词,从而改变概率分布,易被机器学习模型检测,但对人类极难察觉。此外还识别出四个解释类别:时间漂移、美式表达、外语使用及口语化表达。由于检测依赖于此类特征的组合,该分类方法具有鲁棒性,难以被有意伪造者绕过。
原文摘要 · Abstract (English)
We consider the problem of distinguishing human-written creative fiction (excerpts from novels) from similar text generated by an LLM. Our results show that, while human observers perform poorly (near chance levels) on this binary classification task, a variety of machine-learning models achieve accuracy in the range 0.93 - 0.98 over a previously unseen test set, even using only short samples and single-token (unigram) features. We therefore employ an inherently interpretable (linear) classifier (with a test accuracy of 0.98), in order to elucidate the underlying reasons for this high accuracy. In our analysis, we identify specific unigram features indicative of LLM-generated text, one of the most important being that the LLM tends to use a larger variety of synonyms, thereby skewing the probability distributions in a manner that is easy to detect for a machine learning classifier, yet very difficult for a human observer. Four additional explanation categories were also identified, namely, temporal drift, Americanisms, foreign language usage, and colloquialisms. As identification of the AI-generated text depends on a constellation of such features, the classification appears robust, and therefore not easy to circumvent by malicious actors intent on misrepresenting AI-generated text as human work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。