arXiv:2602.07090cs.CRcs.AI2026-02被引 2

针对文本嵌入的隐私泄露问题,提出按概念精准加噪的新防护框架。

Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks

  • 通过可学习掩码识别用户定义概念的敏感维度
  • 采用马氏距离噪声仅扰动敏感维度,隐私泄露降低60%以上
  • 适合需保护特定敏感信息的NLP应用开发者

文本嵌入虽推动了众多自然语言处理应用的发展,却面临嵌入反演攻击带来的严重隐私风险,可能导致敏感属性暴露或原始文本重建。现有差分隐私防御方法假设所有嵌入维度具有相同敏感度,导致噪声过大且性能下降。本文提出SPARSE——一种以用户为中心的概念特异性隐私保护框架。该框架结合(1)可微掩码学习,识别用户定义概念对应的隐私敏感维度;(2)基于维度敏感度校准的马氏机制,施加椭球形噪声。与传统球形噪声注入不同,SPARSE仅对敏感维度进行扰动,保留非敏感语义。在六个数据集、三种嵌入模型及多种攻击场景下评估,SPARSE在持续降低隐私泄露的同时,相比现有最先进差分隐私方法实现了更优的下游任务性能。

原文摘要 · Abstract (English)

Text embeddings enable numerous NLP applications but face severe privacy risks from embedding inversion attacks, which can expose sensitive attributes or reconstruct raw text. Existing differential privacy defenses assume uniform sensitivity across embedding dimensions, leading to excessive noise and degraded utility. We propose SPARSE, a user-centric framework for concept-specific privacy protection in text embeddings. SPARSE combines (1) differentiable mask learning to identify privacy-sensitive dimensions for user-defined concepts, and (2) the Mahalanobis mechanism that applies elliptical noise calibrated by dimension sensitivity. Unlike traditional spherical noise injection, SPARSE selectively perturbs privacy-sensitive dimensions while preserving non-sensitive semantics. Evaluated across six datasets with three embedding models and attack scenarios, SPARSE consistently reduces privacy leakage while achieving superior downstream performance compared to state-of-the-art DP methods.

隐私保护嵌入安全差分隐私文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。