用生成式AI代理模拟用户,大规模审计个性化算法的偏见。
Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

- 用带人格设定的AI代理模拟真实用户行为,实现可控制的黑箱审计。
- 在X平台部署1120个代理,发现算法放大右翼、极化政治内容,且影响因意识形态而异。
- 通过扰动身份信号,揭示算法对不同人群的差异化响应,适合政策研究者和平台安全团队。
个性化算法决定用户在在线平台上看到的内容。审计这些系统困难重重,因为独立审计者仅能通过黑箱访问算法,而个性化依赖于用户的属性、行为及不断演变的交互历史。现有审计方法面临权衡:真实用户研究虽能捕捉真实行为,但成本高且难以控制;而傀儡账户审计虽易扩展,却常依赖预设脚本,缺乏真实性。此外,两者均难以分离用户属性与行为,限制了对个性化的因果理解。为此,我们提出一种基于生成式AI代理的黑箱审计框架,用以构建具有固定人格的合成账号。每个代理基于人口统计和政治调查数据设定人格,通过推理选择行为。由于人格固定,而平台可见信号(如年龄、性别、位置)可实验性扰动,该设计支持对平台如何响应用户属性的反事实分析。作为案例研究,我们在2024年美国大选后不久,在X平台部署了1120个代理,覆盖14种人格和三种反事实条件,收集超20万次内容曝光。结果发现,X的算法推荐流相比时间线流,显著放大有毒、极化、政治化及右翼内容,且放大程度随用户意识形态剧烈变化。反事实分析表明,人口统计信号对内容分发的影响在整体上几乎为零,但在子群体层面方向与强度各异。本研究确立了生成式AI代理作为算法审计的新工具。
原文摘要 · Abstract (English)
Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors have only black-box access to the algorithms, while personalization depends on users' attributes, behavior, and evolving interaction histories. Existing auditing methods face a tradeoff: studies with real users capture realistic behavior but are costly and hard to control, whereas sock-puppet audits scale more easily but often rely on scripted behavior that limits realism. Beyond this, both approaches struggle to decouple user attributes from user behavior, limiting our ability to causally understand personalization. To address this gap, we introduce a framework for black-box audits of personalization algorithms using generative AI agents as behavioral engines for synthetic accounts. Each agent is instantiated with a fixed persona, grounded in demographic and political survey data, and interacts with a platform's content by reasoning about it and choosing actions. Because behavior is fixed within each persona while platform-visible signals such as age, gender, or location can be experimentally perturbed, our design enables counterfactual auditing of how platforms respond to user attributes. As a case study, we deploy 1,120 agents on X shortly after the 2024 U.S. election, spanning 14 personas and three counterfactual conditions, collecting over 200,000 content exposures. We find that X's algorithmic feed amplifies toxic, polarizing, political, and right-leaning content relative to the chronological feed, with amplification varying sharply by user ideology. Counterfactual analyses show that demographic signals affect content delivery in persona-dependent ways: pooled effects are largely null, while subgroup-level effects vary in direction and magnitude. Our work establishes GenAI-based agents as a new tool for algorithmic auditing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。