arXiv:2601.15220cs.CL2026-01ACL被引 2

微调竟会暴露隐私,模型越‘聪明’越不靠谱。

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

  • 用真实数据微调大模型,会悄悄破坏上下文隐私
  • 六种模型五类数据均出现隐私泄露,但性能仍达标
  • 适合关注AI安全、智能体部署的开发者和研究者

我们发现语言模型中一种新现象:对前沿模型进行看似无害的微调,可能导致隐私崩溃。多样且微妙的训练数据模式——如优化帮助性、暴露用户信息、情感对话、调试代码时打印内部变量等——都会削弱上下文隐私保护能力。微调后模型丧失对隐私规范的推理能力,会不当使用工具共享信息,并跨上下文突破记忆边界。隐私崩溃是一种“静默失败”,因为模型在标准安全与效用基准上表现良好,却存在严重隐私漏洞。实验表明,六种模型(闭源与开源)、五类微调数据(真实与受控数据)、两类任务(代理型与记忆型)均出现隐私崩溃。机制分析显示,隐私表征比任务相关特征更易被微调破坏。结果揭示了当前安全评估的关键缺口,尤其在专用智能体部署中。

原文摘要 · Abstract (English)

We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle patterns in training data can degrade contextual privacy, including optimisation for helpfulness, exposure to user information, emotional and subjective dialogue, and debugging code printing internal variables, among others. Fine-tuned models lose their ability to reason about contextual privacy norms, share information inappropriately with tools, and violate memory boundaries across contexts. Privacy collapse is a ``silent failure'' because models maintain high performance on standard safety and utility benchmarks whilst exhibiting severe privacy vulnerabilities. Our experiments show evidence of privacy collapse across six models (closed and open weight), five fine-tuning datasets (real-world and controlled data), and two task categories (agentic and memory-based). Our mechanistic analysis reveals that privacy representations are uniquely fragile to fine-tuning, compared to task-relevant features which are preserved. Our results reveal a critical gap in current safety evaluations, in particular for the deployment of specialised agents.

隐私安全模型微调智能体大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。