小团队协同伪造反馈,可让推荐系统安全机制失效20%。
With a Little Help From My Friends: Collective Manipulation in Risk-Controlling Recommender Systems
- 用1%用户集体虚假标记'不感兴趣',干扰系统风控信号。
- 实测导致普通用户nDCG下降20%,推荐质量严重受损。
- 适合关注算法安全、反作弊的平台方与研究者参考。
推荐系统已成为在线信息分发的核心枢纽,影响用户在各类活动中的行为。为应对这一影响,用户开始组织协作,利用点赞、评论等平台功能引导算法结果,以推广优质内容或抑制有害信息。尽管此类行为可能有益,但也可能被用于恶意操纵,尤其在依赖反馈直接控制风险的推荐系统中。本文研究了这类风险控制型推荐系统的脆弱性:它们通过二值反馈(如“不感兴趣”)结合合意风险控制,可证明地限制用户接触不良内容。我们基于大规模视频平台数据发现,仅需1%的协调用户即可利用系统提供的反馈机制,使非对抗性用户的nDCG下降高达20%。评估了无需了解算法细节的简单攻击策略,结果显示虽然整体推荐质量显著下降,但仅靠报告无法精准压制特定内容组。最后提出一种新策略,将安全保证从群体层面转移至个体层面,实验表明该方法能有效降低协同攻击影响,并保障个人推荐安全性。
原文摘要 · Abstract (English)
Recommendation systems have become central gatekeepers of online information, shaping user behaviour across a wide range of activities. In response, users increasingly organize and coordinate to steer algorithmic outcomes toward diverse goals, such as promoting relevant content or limiting harmful material, relying on platform affordances -- such as likes, reviews, or ratings. While these mechanisms can serve beneficial purposes, they can also be leveraged for adversarial manipulation, particularly in systems where such feedback directly informs safety guarantees. In this paper, we study this vulnerability in recently proposed risk-controlling recommender systems, which use binary user feedback (e.g., "Not Interested") to provably limit exposure to unwanted content via conformal risk control. We empirically demonstrate that their reliance on aggregate feedback signals makes them inherently susceptible to coordinated adversarial user behaviour. Using data from a large-scale online video-sharing platform, we show that a small coordinated group (comprising only 1% of the user population) can induce up to a 20% degradation in nDCG for non-adversarial users by exploiting the affordances provided by risk-controlling recommender systems. We evaluate simple, realistic attack strategies that require little to no knowledge of the underlying recommendation algorithm and find that, while coordinated users can significantly harm overall recommendation quality, they cannot selectively suppress specific content groups through reporting alone. Finally, we propose a mitigation strategy that shifts guarantees from the group level to the user level, showing empirically how it can reduce the impact of adversarial coordinated behaviour while ensuring personalized safety for individuals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。