arXiv:2608.18610cs.LGcs.AI2026-08中稿 · IEEE ICDM 2026

提出新攻击方法,揭露加噪文本嵌入的隐私漏洞

Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings

论文配图:Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings
图 1 · 摘自论文原文
  • 设计去噪感知反演框架,从噪声嵌入中重建原文
  • 在多个指标上提升32%-154%,超越现有方法
  • 警示仅靠高斯噪声防护已不足,适合隐私研究者

稠密文本嵌入广泛用于数据挖掘、信息检索和下游机器学习系统,但近期嵌入反演攻击表明其可能暴露原始文本的大量信息,引发严重隐私泄露风险。常见防御是添加高斯噪声以释放扰动后的嵌入,该方法简单有效,且对下游任务性能影响较小。然而,目前尚不清楚此类加噪嵌入是否能抵御明确考虑扰动过程的自适应攻击者。本文研究在噪声保护设置下的文本嵌入反演问题,即攻击者只能观测到噪声嵌入,无法获取干净嵌入目标。我们首先分析现有生成式反演方法为何在此场景下失效,并识别出‘双重噪声陷阱’这一根本障碍。为此,我们提出DAEI,一种结合残差去噪自编码器与生成式文本反演的管道,其中去噪器通过无监督方式使用Stein无偏风险估计进行训练,仅凭噪声观测即可实现去噪。大量实验表明,DAEI在BLEU上相对基线提升约154%,在词级F1和ROUGE-L上分别提升32%–60%。该优异的反演性能挑战了当前认为简单高斯扰动足以防止敏感信息泄露的普遍假设。

原文摘要 · Abstract (English)

Dense text embeddings are widely used in data mining, retrieval, and downstream machine learning systems due to their compact and semantically rich representations, but recent embedding inversion attacks have shown that they can expose substantial information about the original text, leading to serious privacy leakage risks. A common defense is to release perturbed embeddings by adding Gaussian noise, which is simple yet effective against standard inversion attacks and does not significantly degrade embedding utility for downstream tasks. However, it remains unclear whether such noise-protected embeddings are sufficiently safe against adaptive attackers that explicitly account for the perturbation process. In this paper, we study text embedding inversion in a noise-protected setting, where the attacker can observe only noisy embeddings and has no access to clean embedding targets. We first analyze why existing generative inversion methods fail under this setting and identify a "Double Noise Trap", which fundamentally prevents standard generative inversion models from achieving high-quality reconstruction. To address this challenge, we propose DAEI, a denoising-aware embedding inversion pipeline that combines a residual denoising autoencoder with generative text inversion where the denoiser is trained in an unsupervised manner using Stein's unbiased risk estimate to enable denoising from noisy observations alone. Extensive experiments show that DAEI achieves approximately 154\% relative improvement in BLEU over the existing generative inversion baseline, while also improving token-level F1 and ROUGE-L by 32--60\%. The promising inversion performance of DAEI challenges the prevailing assumption that simple Gaussian perturbation is sufficient to prevent sensitive information leakage from embedding representations.

隐私安全嵌入反演去噪文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。