arXiv:2608.23382cs.LGcs.CR2026-08

提出更紧的隐私编码不可逆性边界,支持确定性和随机编码。

Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding

论文配图:Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding
图 1 · 摘自论文原文
  • 基于编码器谱结构设计新边界,突破原有限制。
  • 在多种编码器和攻击下边界均成立,且比旧方法更紧。
  • 适合关注数据隐私理论保障的研究者与工程师。

实例编码是一种常见的隐私增强技术,通过编码过程在共享敏感数据前将其转换,以期在保留数据效用的同时难以还原原始信息。然而,多数研究缺乏对编码过程实际不可逆性的理论保证。近期工作推导出均方误差(MSE)上界,为该领域首个理论成果之一,但存在三方面局限:上界常过于宽松,仅适用于随机编码器(排除了大量实践中使用的确定性编码器),且仅限于MSE度量。本文提出一类新边界,可(1)更紧致,(2)适用于完全确定性编码器,(3)扩展至其他基于范数的相似性度量,关键在于合理建模编码器的谱结构。我们在多种编码器、数据集和攻击场景下评估新边界,结果表明其始终成立,且显著优于现有方法。

原文摘要 · Abstract (English)

Instance encoding is a popular empirical technique for privacy enhancement when sharing data to an untrusted server. It transforms sensitive data through an encoding process before sharing, with the hope that the encoding process retains utility but makes it hard to reconstruct the original data. However, most work offers no theoretical guarantee that the encoding process is actually irreversible. A recent work derived a mean-squared error (MSE) bound limiting any adversary's reconstruction accuracy, offering one of the first theoretical results in this domain. This bound, however, has three critical limitations: it is often too loose, only works with randomized encoders (excluding many deterministic encoders practitioners use), and only bounds MSE. We introduce a family of new bounds that (1) are tighter, (2) applicable even to fully deterministic encoders, and (3) can extend beyond MSE to other norm-based similarity metrics, by properly accounting for the encoder's spectral structure. We evaluate our bounds across a range of encoders, datasets, and attacks, showing they hold consistently and improve upon the existing bound.

隐私编码理论保障谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。