arXiv:2608.22161cs.AIcs.CL2026-08

让多篇生成文本协同伪装,防身份聚合追踪。

Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification

论文配图:Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification
图 1 · 摘自论文原文
  • 在文本合集层面联合生成,避免单篇优化导致的隐私泄露
  • 随文本数量增加,账号关联性显著降低,跨领域攻击下仍有效
  • 适合需保护多篇发文身份的用户,如社交平台匿名发布者

在线用户常以同一身份发布多篇文本,攻击者可借此构建完整作者画像。现有隐私保护方法独立优化每篇文本,忽视跨文档关联带来的聚合风险。本文提出聚合感知的合成文本生成框架AAST,通过在文本集合层面联合选择生成内容,抵御溯源与验证攻击,包括攻击样本来自未参与生成的文本领域的跨领域场景。在同领域、跨领域、神经与非神经风格特征攻击下,实验表明随着文本集合规模增大,账号级关联性持续下降,同时保持语义质量、语言自然度和情感一致性。

原文摘要 · Abstract (English)

Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation methods optimize privacy independently for each document, leaving them blind to cross-document correlations that make aggregation dangerous. We propose Aggregation-Aware Synthetic Text Generation (AAST), a framework that addresses this gap by jointly selecting synthetic texts at the bundle level rather than optimizing each text in isolation. AAST targets attribution and verification attacks, including cross-genre settings where attacker references come from a genre not observed during generation or selection. Experiments across same-genre, cross-genre, neural, and independent non-neural stylometric attacks show that AAST lowers account-level linkability as bundle size grows, while preserving semantic quality, linguistic acceptability, and sentiment alignment.

文本生成隐私保护身份隐藏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。