arXiv:2506.23855cs.CRcs.AI2025-06KDD被引 1

生成可保护隐私的模拟广告数据,用于真实研究广告接口的隐私风险。

Differentially Private Synthetic Data Release for Topics API Outputs

  • 基于差分隐私计算真实数据的统计特征,构建可模拟接口输出的生成模型。
  • 合成数据在重识别风险上与真实数据高度一致,且提供理论化隐私保障。
  • 开源数据集助力学术界深入分析隐私保护广告接口,推动透明度提升。

隐私保护广告API的隐私特性分析是学术界、产业界和监管机构关注的重点。然而,由于缺乏公开可用的数据,相关实证研究面临困难。可靠的隐私分析需要包含真实API输出的大型数据集,但隐私顾虑限制了此类数据的公开。本文提出一种新方法,生成既具备足够真实性以支持准确研究,又提供强隐私保护的合成数据。聚焦谷歌Chrome隐私沙盒中的主题API(Topics API),我们构建了一个差分隐私数据集,其重识别风险特性与真实数据高度匹配。通过首先计算大量差分隐私统计量来描述API输出随时间演变的规律,再设计参数化序列分布并优化其参数以拟合这些统计量,最终从该分布中采样生成合成数据。本工作还开源了经匿名处理的数据集,旨在支持外部研究人员进行深度分析,并复现或开展未来研究。我们相信该工作将促进对隐私保护广告接口隐私特性的透明化理解。

原文摘要 · Abstract (English)

The analysis of the privacy properties of Privacy-Preserving Ads APIs is an area of research that has received strong interest from academics, industry, and regulators. Despite this interest, the empirical study of these methods is hindered by the lack of publicly available data. Reliable empirical analysis of the privacy properties of an API, in fact, requires access to a dataset consisting of realistic API outputs; however, privacy concerns prevent the general release of such data to the public. In this work, we develop a novel methodology to construct synthetic API outputs that are simultaneously realistic enough to enable accurate study and provide strong privacy protections. We focus on one Privacy-Preserving Ads APIs: the Topics API, part of Google Chrome's Privacy Sandbox. We developed a methodology to generate a differentially-private dataset that closely matches the re-identification risk properties of the real Topics API data. The use of differential privacy provides strong theoretical bounds on the leakage of private user information from this release. Our methodology is based on first computing a large number of differentially-private statistics describing how output API traces evolve over time. Then, we design a parameterized distribution over sequences of API traces and optimize its parameters so that they closely match the statistics obtained. Finally, we create the synthetic data by drawing from this distribution. Our work is complemented by an open-source release of the anonymized dataset obtained by this methodology. We hope this will enable external researchers to analyze the API in-depth and replicate prior and future work on a realistic large-scale dataset. We believe that this work will contribute to fostering transparency regarding the privacy properties of Privacy-Preserving Ads APIs.

隐私保护合成数据差分隐私广告系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。