arXiv:2607.01824cs.LG2026-07

揭露众包事实核查中伪造共识的漏洞,揭示少数人可低成本操控结果

Gaming Consensus: Coordinated Manipulation in Crowdsourced Fact-Checking

论文配图:Gaming Consensus: Coordinated Manipulation in Crowdsourced Fact-Checking
图 1 · 摘自论文原文
  • 利用矩阵分解的隐式表示,设计策略性投票伪造跨群体支持
  • 实测仅需少于10次评分即可让10.7%低质量标注通过共识阈值
  • 发现'无帮助'评分反而提升得分,适合平台安全与算法设计者阅读

众包事实核查系统被X、Meta、TikTok和Google等主流社交媒体采用,旨在不依赖中心化编辑的情况下大规模遏制误导信息。这些系统基于一种核心机制:当标注获得不同观点人群的支持时才被视为可信,而非简单多数支持。目前公开披露的桥梁算法均基于矩阵分解,如X和Meta所用,并附加反滥用与对抗组织操纵的组件。本文聚焦该机制的核心矩阵分解部分,从理论与实证角度分析协同用户如何利用隐向量策略性投票,制造虚假共识。基于历史生产数据,我们发现仅需不到10次评分,即可使高达10.7%的低质量标注突破共识阈值。理论分析进一步揭示反直觉现象:将标注评为‘无帮助’可能反而提高其有用性得分,并建立量化操纵成本的模型。相关缓解措施已部署至X的Community Notes算法中。

原文摘要 · Abstract (English)

Crowdsourced fact-checking systems have been adopted by major social media companies such as X, Meta, TikTok and Google with the aim of combating misleading information at scale without relying on centralized editorial control. These systems have been developed around a common underlying concept: a bridging mechanism that identifies notes flagging misleading information when they receive support from people with different perspectives rather than simple majority support. To our knowledge the only publicly disclosed bridging algorithms deployed for fact-checking are based on matrix factorization, as deployed by both X and Meta, augmented with additional components addressing abuse, targeted manipulation, and contributor brigades. This work examines the core matrix factorization portion of these systems, presenting theoretical and empirical evaluations of the degree to which coordinated users could vote strategically by leveraging the latent representations to fabricate the appearance of synthetic consensus within the bridging mechanism. Using historic production data, we find that up to 10.7% of lower quality notes could be manipulated above consensus thresholds using less than 10 ratings. We complement these findings with a theoretical analysis, revealing counterintuitively that rating a note as "Not Helpful" can increase its helpfulness score, as well as a cost model quantifying manipulation effort. We have developed and deployed mitigations within X's Community Notes algorithm to address synthetic consensus.

众包核查共识操纵矩阵分解安全机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。