AI可伪装作者风格,让身份验证失效
Masks and Mimicry: Strategic Obfuscation and Impersonation Attacks on Authorship Verification
- 用LLM改写文本,隐藏原作者风格
- 攻击成功率最高达92%(隐藏)和78%(模仿)
- 警示学术与法律领域需防范文本伪造
大型语言模型(LLMs)虽提升了文档作者身份识别的准确性,但也为恶意攻击者提供了新手段。本文评估了作者身份验证模型在对抗性攻击下的鲁棒性,重点研究两类攻击:无目标攻击(作者风格隐蔽)和有目标攻击(作者风格模仿)。通过扰动原文本,在保持语义不变的前提下,成功使作者验证模型误判,实现最高92%的隐蔽攻击成功率和78%的模仿攻击成功率。
原文摘要 · Abstract (English)
The increasing use of Artificial Intelligence (AI) technologies, such as Large Language Models (LLMs) has led to nontrivial improvements in various tasks, including accurate authorship identification of documents. However, while LLMs improve such defense techniques, they also simultaneously provide a vehicle for malicious actors to launch new attack vectors. To combat this security risk, we evaluate the adversarial robustness of authorship models (specifically an authorship verification model) to potent LLM-based attacks. These attacks include untargeted methods - \textit{authorship obfuscation} and targeted methods - \textit{authorship impersonation}. For both attacks, the objective is to mask or mimic the writing style of an author while preserving the original texts' semantics, respectively. Thus, we perturb an accurate authorship verification model, and achieve maximum attack success rates of 92\% and 78\% for both obfuscation and impersonation attacks, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。