通过内在分布指纹,检测大模型微调中的数据侵权行为。
Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints

- 基于语义与词法的持久交集,提取数据指纹
- 在医疗法律任务中优于基线,对抗改写与蒸馏攻击
- 适合数据版权保护者和安全研究人员使用
定制化大语言模型的普及带来了未经授权使用专有数据进行微调的数据知识产权(Data IP)侵权风险。现有审计技术受限于训练过程中的干预要求,且在恶意混淆(如数据改写、知识蒸馏)下表现脆弱。本文提出后置式审计框架DPA(Distribution Provenance Audit),其核心洞察是:无论采用何种规避策略,为维持模型性能,必须保留语义实质与词汇形式的基本交集。DPA将这一恒定交集作为内在分布指纹,通过无偏输出采样构建统计假设检验,可靠拒绝非使用假设。在医疗与法律微调任务上的实验表明,DPA持续优于现有基线,对采用改写和蒸馏的对抗性训练者仍具鲁棒性。此外,我们揭示了双重用途矛盾:高保真分布指纹既能实现可靠审计,也可能被用于隐私攻击。
原文摘要 · Abstract (English)
The proliferation of customized Large Language Models (LLMs) poses critical risks of Data Intellectual Property (Data IP) infringement via unauthorized fine-tuning on proprietary data. Existing audit techniques are limited, as they require intervention during data preparation or training and remain fragile under malicious obfuscations such as data paraphrasing and knowledge distillation. We propose \textit{Distribution Provenance Audit (DPA)}, a post-hoc framework for auditing data IP infringement in LLM fine-tuning under black-box and malicious settings. DPA is grounded in a critical insight: regardless of fine-tuning tactics to evade provenance, the practical necessity of maintaining utility constrains the model to preserve the fundamental intersection of semantic substance and lexical form. Accordingly, DPA captures this persistent lexical-semantic intersection as intrinsic distributional fingerprints. The framework formulates the audit as a statistical hypothesis test, effectively quantifying these fingerprints via unbiased output sampling to reliably reject the null hypothesis of non-usage. Extensive experiments on medical and legal fine-tuning tasks show that DPA consistently outperforms existing baselines, remaining robust against adversarial trainers employing paraphrasing and knowledge distillation. We further highlight a fundamental dual-use tension: the same high-fidelity distributional fingerprints enabling reliable auditing may also facilitate privacy attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。