arXiv:2601.12937cs.CRcs.AI2026-01

用语义不变的改写数据,让版权审计攻击失效。

On the Evidentiary Limits of Membership Inference for Copyright Auditing

  • 用稀疏自编码器引导重写训练数据,保持语义但改变词汇结构。
  • 主流版权推断攻击在改写后模型上准确率大幅下降。
  • 适合关注大模型版权审计可靠性的研究者与法律从业者。

随着大语言模型(LLMs)训练数据日益不透明,成员推理攻击(MIAs)被提出用于审计训练过程中是否使用了受版权保护的文本。然而,在现实条件下其可靠性受到质疑。本文探讨在对抗性版权纠纷中,当被指控的模型开发者可能隐藏训练数据但保留语义内容时,MIAs能否作为可采信证据,并通过法官-检察官-被告通信协议形式化该场景。为测试鲁棒性,我们提出SAGE(结构感知的SAE引导提取)框架,利用稀疏自编码器(SAEs)生成语义不变但词汇结构变化的改写数据。实验表明,经过SAGE改写数据微调后的模型,使现有顶尖MIAs性能显著下降,说明其信号对语义保持型变换不鲁棒。尽管某些微调场景仍存在少量泄露,结果表明MIAs在对抗环境下脆弱,无法单独作为大模型版权审计的可靠手段。

原文摘要 · Abstract (English)

As large language models (LLMs) are trained on increasingly opaque corpora, membership inference attacks (MIAs) have been proposed to audit whether copyrighted texts were used during training, despite growing concerns about their reliability under realistic conditions. We ask whether MIAs can serve as admissible evidence in adversarial copyright disputes where an accused model developer may obfuscate training data while preserving semantic content, and formalize this setting through a judge-prosecutor-accused communication protocol. To test robustness under this protocol, we introduce SAGE (Structure-Aware SAE-Guided Extraction), a paraphrasing framework guided by Sparse Autoencoders (SAEs) that rewrites training data to alter lexical structure while preserving semantic content and downstream utility. Our experiments show that state-of-the-art MIAs degrade when models are fine-tuned on SAGE-generated paraphrases, indicating that their signals are not robust to semantics-preserving transformations. While some leakage remains in certain fine-tuning regimes, these results suggest that MIAs are brittle in adversarial settings and insufficient, on their own, as a standalone mechanism for copyright auditing of LLMs.

版权审计成员推理语义保持大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。