arXiv:2603.29454cs.CL2026-03

用大模型模仿他人写作,仍难逃文本溯源检测。

Authorship Impersonation via LLM Prompting does not Evade Authorship Verification Methods

  • 用GPT-4o在四种提示下生成邮件、短信等文本进行模仿
  • 所有模仿文本均被现有系统准确识别,部分方法反向更准
  • 大模型文本词汇多样性高,反而暴露了伪造痕迹

作者身份验证(AV)是法医语言学中的关键任务,旨在判断某段文本是否由特定个体所写。尽管历史上人为模仿作者风格长期存在威胁,但大型语言模型(LLMs)的兴起带来了新挑战:攻击者可能利用这些工具模仿他人写作风格。本研究探究了提示后的LLM能否生成逼真的模仿文本,并规避现有法医AV系统。以GPT-4o为攻击模型,在邮件、短信、社交媒体三种文体下,设置四种提示条件生成模仿文本,并在似然比框架下评估其对非神经网络方法(n-gram tracing、Ranking-Based Impostors Method、LambdaG)和神经方法(AdHominem、LUAR、STAR)的欺骗能力。结果表明,LLM生成文本无法充分复现作者个体特征,未能绕过现有AV系统。甚至有部分方法在识别模仿文本时的准确率高于识别真实负样本。整体说明,尽管LLM易得,当前AV系统对多文体下的初级模仿仍具鲁棒性。进一步分析发现,这种反直觉的稳健性部分源于LLM生成文本固有的更高词汇多样性与熵值。

原文摘要 · Abstract (English)

Authorship verification (AV), the task of determining whether a questioned text was written by a specific individual, is a critical part of forensic linguistics. While manual authorial impersonation by perpetrators has long been a recognized threat in historical forensic cases, recent advances in large language models (LLMs) raise new challenges, as adversaries may exploit these tools to impersonate another's writing. This study investigates whether prompted LLMs can generate convincing authorial impersonations and whether such outputs can evade existing forensic AV systems. Using GPT-4o as the adversary model, we generated impersonation texts under four prompting conditions across three genres: emails, text messages, and social media posts. We then evaluated these outputs against both non-neural AV methods (n-gram tracing, Ranking-Based Impostors Method, LambdaG) and neural approaches (AdHominem, LUAR, STAR) within a likelihood-ratio framework. Results show that LLM-generated texts failed to sufficiently replicate authorial individuality to bypass established AV systems. We also observed that some methods achieved even higher accuracy when rejecting impersonation texts compared to genuine negative samples. Overall, these findings indicate that, despite the accessibility of LLMs, current AV systems remain robust against entry-level impersonation attempts across multiple genres. Furthermore, we demonstrate that this counter-intuitive resilience stems, at least in part, from the higher lexical diversity and entropy inherent in LLM-generated texts.

作者识别大模型安全文本溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。