用大模型识别作者风格,能准确区分不同作家的文本。
A Stylometric Application of Large Language Models
- 为每位作者单独训练GPT-2,学习其独特写作风格。
- 同一作者的未见文本预测准确率显著高于其他作者。
- 成功验证了《奥兹国》第15册真实作者为R.P.汤普森。
我们证明大型语言模型(LLMs)可用于区分不同作者的写作风格。具体而言,一个从零开始在单个作者作品上训练的GPT-2模型,对同一作者的未见文本预测准确率,显著高于对其他作者文本的预测。这表明,基于某位作者作品训练的模型,会内化其独特的写作风格。我们首先在八位已知作者的小说上验证该方法。此外,该方法还用于确认《奥兹国》系列第15本书的真实作者为R. P. Thompson,而非原署名的F. L. Baum。
原文摘要 · Abstract (English)
We show that large language models (LLMs) can be used to distinguish the writings of different authors. Specifically, an individual GPT-2 model, trained from scratch on the works of one author, will predict held-out text from that author more accurately than held-out text from other authors. We suggest that, in this way, a model trained on one author's works embodies the unique writing style of that author. We first demonstrate our approach on books written by eight different (known) authors. We also use this approach to confirm R. P. Thompson's authorship of the well-studied 15th book of the Oz series, originally attributed to F. L. Baum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。