arXiv:2502.08933cs.LG2025-02被引 2

用强化学习模拟用户行为,自动检测社交平台推荐内容风险。

AutoLike: Auditing Social Media Recommendations through User Interactions

  • 通过强化学习模拟用户点赞,诱导推荐系统暴露内容偏好。
  • 在TikTok上验证对9类话题的识别能力,8项实验成功引导推荐方向。
  • 适合监管机构用于审查算法推荐中的有害内容风险。

现代社交媒体平台(如TikTok、Facebook、YouTube)依赖推荐系统根据用户互动为用户提供个性化内容,例如“为你推荐”页面。然而,这些复杂算法可能无意中推送与自残、心理健康和饮食失调相关的问题内容。我们提出AutoLike框架,用于审计社交平台推荐系统中特定主题及其情感倾向。为实现自动化,我们将问题建模为强化学习任务:通过模拟用户交互(如点赞)驱动推荐系统呈现特定类型内容。以TikTok为例,我们评估了AutoLike在9个兴趣主题上的自动识别能力,并开展8项实验,验证其引导推荐系统向特定主题和情感偏移的有效性。该方法可协助监管机构审计推荐系统中的有害内容风险。(警告:本文包含可能令人不适或有害的定性示例。)

原文摘要 · Abstract (English)

Modern social media platforms, such as TikTok, Facebook, and YouTube, rely on recommendation systems to personalize content for users based on user interactions with endless streams of content, such as "For You" pages. However, these complex algorithms can inadvertently deliver problematic content related to self-harm, mental health, and eating disorders. We introduce AutoLike, a framework to audit recommendation systems in social media platforms for topics of interest and their sentiments. To automate the process, we formulate the problem as a reinforcement learning problem. AutoLike drives the recommendation system to serve a particular type of content through interactions (e.g., liking). We apply the AutoLike framework to the TikTok platform as a case study. We evaluate how well AutoLike identifies TikTok content automatically across nine topics of interest; and conduct eight experiments to demonstrate how well it drives TikTok's recommendation system towards particular topics and sentiments. AutoLike has the potential to assist regulators in auditing recommendation systems for problematic content. (Warning: This paper contains qualitative examples that may be viewed as offensive or harmful.)

推荐系统算法审计强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。