小规模私有数据微调模型可被用来重构特定用户隐私信息。
Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models

- 设计针对性攻击方法COVA,基于上下文覆盖优化解码过程
- 少量上下文知识即可显著提升隐私信息重建成功率
- 医疗与法律领域数据验证了微调模型存在真实隐私泄露风险
监督微调(SFT)已成为将具备丰富预训练知识的大语言模型适配到特定领域指令跟随任务的主要方法。SFT数据集由指令-响应对构成,常包含用户提供的敏感信息,如个人身份信息(PII),引发隐私担忧。本文研究针对特定身份的PII从私有SFT数据微调模型中重构的问题。我们在医疗和法律等敏感领域构建了多轮、以用户为中心的问答数据集,引入PII以实现对泄露问题的真实评估。提出一种覆盖率感知的解码算法COVA,用于在前缀攻击下重构目标PII。通过COVA,我们分析了关于目标用户的信息量如何影响其PII的恢复。在多个数据集上,COVA持续优于基线解码方法;即使仅有有限的上下文知识,也能大幅提高攻击成功率。结果表明,小型私有SFT数据集可在微调过程中学习到目标PII关联,从而导致显著的隐私泄露。
原文摘要 · Abstract (English)
Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to domain-specific, instruction-following tasks. SFT datasets, composed of instruction-response pairs, often include user-provided information that may contain sensitive data such as personally identifiable information (PII), raising privacy concerns. This paper studies the problem of targeted PII reconstruction from models fine- tuned on proprietary SFT data, in which an adversary attempts to recover PII associated with a specific identity. We construct multi-turn, user-centric Q&A datasets in sensitive domains, specifically medical and legal settings, that incorporate PII to enable realistic evaluation of leakage. We then propose COVA, a coverage-aware decoding algorithm for targeted PII reconstruction under prefix-based attacks. Using COVA, we study how the amount of information available about a target user affects the recovery of their PII from SFT models. Across datasets, COVA consistently improves PII reconstruction over baseline decoding methods, and even limited contextual knowledge can substantially increase an adversary's success. Our findings demonstrate that small, proprietary SFT datasets can induce meaningful privacy leakage through the reconstruction of target-PII associations learned during fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。