arXiv:2509.18874cs.HCcs.AI2025-09被引 1

用大模型从广告曝光中逆向推断用户隐私属性,揭示网络广告的隐蔽风险。

When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs

  • 用零样本多模态大模型作为攻击引擎,从广告内容推断用户隐私信息。
  • 在43.5万条广告数据上,准确率超普查基准,成本仅为人工的1/223。
  • 短时间观察即可完成精准画像,说明无需长期追踪也能实现攻击。

监管对显式定位的限制并未消除网页上的算法画像,因优化系统仍会基于用户的私有属性调整广告投放。强大且无需训练的零样本多模态大语言模型(LLMs)显著降低了利用这些隐含信号进行对抗性推断的门槛。我们研究了这一新兴社会风险:攻击者如何仅通过被动观看广告,就利用这些信号反向推导出用户的私有属性。本文提出一种新流程,以大模型作为对抗性推断引擎,实现自然语言层面的画像分析。在涵盖891名用户、超过43.5万条Facebook广告曝光的纵向数据集上开展大规模研究,评估从被动在线广告观察中推断私有属性的可行性与精度。结果表明,现成的大模型能准确重建复杂用户属性,如政治倾向、就业状态和教育水平,持续优于强普查基线,且在成本(低223倍)和时间(快52倍)上远超人类判断。关键发现是:即使在极短观察窗口内,也可实现可操作的画像,表明长期追踪并非成功攻击的必要条件。这些结果首次提供了实证证据:广告流构成高保真数字足迹,支持跨平台画像,天然规避当前平台防护机制,暴露出广告生态系统的系统性漏洞,凸显生成式AI时代负责任网络人工智能治理的紧迫性。代码已公开于 https://github.com/Breezelled/when-ads-become-profiles。

原文摘要 · Abstract (English)

Regulatory limits on explicit targeting have not eliminated algorithmic profiling on the Web, as optimisation systems still adapt ad delivery to users' private attributes. The widespread availability of powerful zero-shot multimodal Large Language Models (LLMs) has dramatically lowered the barrier for exploiting these latent signals for adversarial inference. We investigate this emerging societal risk, specifically how adversaries can now exploit these signals to reverse-engineer private attributes from ad exposure alone. We introduce a novel pipeline that leverages LLMs as adversarial inference engines to perform natural language profiling. Applying this method to a longitudinal dataset comprising over 435,000 Facebook ad impressions collected from 891 users, we conducted a large-scale study to assess the feasibility and precision of inferring private attributes from passive online ad observations. Our results demonstrate that off-the-shelf LLMs can accurately reconstruct complex user private attributes, including party preference, employment status, and education level, consistently outperforming strong census-based priors and matching or exceeding human social perception at only a fraction of the cost (223x lower) and time (52x faster) required by humans. Critically, actionable profiling is feasible even within short observation windows, indicating that prolonged tracking is not a prerequisite for a successful attack. These findings provide the first empirical evidence that ad streams serve as a high-fidelity digital footprint, enabling off-platform profiling that inherently bypasses current platform safeguards, highlighting a systemic vulnerability in the ad ecosystem and the urgent need for responsible web AI governance in the generative AI era. The code is available at https://github.com/Breezelled/when-ads-become-profiles.

隐私安全大模型攻击广告画像生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。