提出隐私保护的注意力机制,首次分析了上下文学习的隐私-精度权衡。
How Private is Your Attention? Bridging Privacy with In-Context Learning
- 设计差分隐私预训练算法,保护线性注意力头的隐私
- 揭示优化与隐私噪声间的根本矛盾,量化隐私代价
- 对对抗性提示扰动鲁棒,适合高敏感场景应用
上下文学习(ICL)——即基于Transformer的模型在推理时通过提供示例完成新任务的能力——已成为现代语言模型的重要特征。尽管已有研究探讨了ICL的机制,但其在正式隐私约束下的可行性仍缺乏深入探索。本文提出一种针对线性注意力头的差分隐私预训练算法,并首次对线性回归中的上下文学习进行了隐私-准确率权衡的理论分析。结果揭示了优化过程与隐私引入噪声之间的根本张力,形式化捕捉了迭代训练中观察到的行为。此外,我们证明该方法对训练提示的对抗性扰动具有鲁棒性,而标准岭回归则不具备此特性。所有理论结论均通过多样设置下的大量模拟实验验证。
原文摘要 · Abstract (English)
In-context learning (ICL)-the ability of transformer-based models to perform new tasks from examples provided at inference time-has emerged as a hallmark of modern language models. While recent works have investigated the mechanisms underlying ICL, its feasibility under formal privacy constraints remains largely unexplored. In this paper, we propose a differentially private pretraining algorithm for linear attention heads and present the first theoretical analysis of the privacy-accuracy trade-off for ICL in linear regression. Our results characterize the fundamental tension between optimization and privacy-induced noise, formally capturing behaviors observed in private training via iterative methods. Additionally, we show that our method is robust to adversarial perturbations of training prompts, unlike standard ridge regression. All theoretical findings are supported by extensive simulations across diverse settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。