研究互动广告中用户属性如何被推断,揭示隐私泄露风险与防御策略。
Attribute Inference from Interactive Targeted Ads
- 构建模拟器分离投放、互动与披露环节,建模隐私泄露路径。
- 160次投放后,贝叶斯与监督攻击达0.64~0.65的AUC,可有效推断属性。
- 聚合报告与随机披露是强防御手段,适合关注广告隐私的研究者。
目标广告系统可将广告主选定的受众与展示广告单元匹配,并暴露用户可见行为。当互动仍与触发它的广告活动关联时,广告主可能获得与特定用户相关的观测结果,而非仅限于汇总报告。本文将该通道建模为属性推断的噪声预言机,分离了定位条件、曝光、互动与披露环节,捕捉了资格与交付之间、互动与广告主可见性之间的差距。我们基于公开数据校准的合成人群构建可复现的基准测试,每组人群均含已知敏感标签。生成的广告语义层提供主题变体和响应先验。模拟器生成真实值、事件轨迹、披露观测及评估指标。评估在常见广告与披露定义下比较了贝叶斯、监督、正负样本及自适应攻击。最终评估使用四种主题变体、七个模拟种子和两种互动设置。重复投放且身份暴露条件下产生可观测但有限的推断信号。在主设置下,贝叶斯与监督攻击在160次投放后达到约0.64 AUC,更高互动设置下可达0.65。披露政策是最强控制手段;聚合报告可消除与用户绑定的输入。类型过滤与随机披露可降低释放信号。本研究提供模型、工具与防御评估方法,用于交互式定向广告中的隐私分析。代码已开源:https://github.com/P-HOW/Interactive-Ad-Oracle。
原文摘要 · Abstract (English)
Targeted advertising systems can pair audiences selected by advertisers with ad units that expose visible user actions. When an interaction remains linked to the campaign that elicited it, the advertiser may receive an observation tied to a user rather than only an aggregate report. We model that channel as a noisy oracle for attribute inference. The model separates targeting predicates, exposure, interaction, and disclosure. These boundaries capture the gap between eligibility and delivery, and the gap between interaction and advertiser visibility. We build a reproducible benchmark using synthetic populations calibrated with public data, each with known sensitive labels. A generated campaign semantics layer provides topic variants and response priors. The simulator generates the ground truth, event traces, disclosed observations, and metrics. The evaluation compares Bayesian, supervised, positive and unlabeled, and adaptive attacks under common campaign and disclosure definitions. The final evaluation uses four topic variants, seven simulator seeds, and two interaction settings. Repeated campaigns with identity exposure produce measurable but bounded inference signal. At $160$ campaigns, Bayesian and supervised attacks reach about $0.64$ AUC in the main setting and about $0.65$ AUC in the higher interaction setting. Disclosure policy is the strongest control. Aggregate reporting removes the evaluated oracle input tied to users. Type filtering and randomized disclosure reduce the released signal. The result is a model, artifact, and defense evaluation method for privacy in interactive targeted advertising. The code is available at https://github.com/P-HOW/Interactive-Ad-Oracle.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。