arXiv:2601.11128cs.SIcs.HC2026-01AAAI被引 2

构建3900万条内容治理数据集,支持系统性研究平台干预效果

The Big Ban Theory: A Pre- and Post-Intervention Dataset of Online Content Moderation Actions

  • 收集Reddit和Voat上25类干预措施前后的用户行为数据
  • 覆盖超33.9万用户与近3900万条帖子,含干预前后三个月数据
  • 提供标准化元数据,助力可复现的治理效果对比研究

在线平台依赖内容治理干预来遏制仇恨言论、攻击性内容及虚假信息传播。然而,现有研究受限于单一或少数干预案例,缺乏全面数据支持,导致研究者需重复采集数据,难以开展系统性比较。为此,我们推出《大爆炸理论》(TBBT)——一个涵盖25种不同类型、严重程度与范围的干预措施的大规模数据集,包含来自Reddit和Voat的33.9万以上用户及近3900万条发布内容。每个干预事件均提供干预前3个月与后3个月的标准化元数据及去标识化用户活动记录,支持一致、可比的分析。此外,我们还提供了数据描述性探索分析及多个研究使用案例,旨在推动内容治理干预效果研究走向更系统、可复现与可比较的方向。

原文摘要 · Abstract (English)

Online platforms rely on moderation interventions to curb harmful behavior such as hate speech, toxicity, and the spread of mis- and disinformation. Yet research on the effects and possible biases of such interventions faces multiple limitations. For example, existing works frequently focus on single or a few interventions, due to the absence of comprehensive datasets. As a result, researchers must typically collect the necessary data for each new study, which limits opportunities for systematic comparisons. To overcome these challenges, we introduce The Big Ban Theory (TBBT) -- a large dataset of moderation interventions. TBBT covers 25 interventions of varying type, severity, and scope, comprising in total over 339K users and nearly 39M posted messages on Reddit and Voat. For each intervention, we provide standardized metadata and pseudonymized user activity collected three months before and after its enforcement, enabling consistent and comparable analyses of intervention effects. In addition, we provide a descriptive exploratory analysis of the dataset, along with several use cases of how it can support research on content moderation. With this dataset, we aim to support researchers studying the effects of moderation interventions and to promote more systematic, reproducible, and comparable research.

内容治理数据集社交媒体干预分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。