arXiv:2506.10245cs.CLcs.AI2025-06被引 2

构建首个葡萄牙语细粒度仇恨言论数据集,支持九类少数群体检测。

ToxSyn-PT: A Synthetic Fine-Grained Dataset of Minority-Targeted Toxic Language in Portuguese

  • 通过四阶段可控生成法构造数据,含讽刺、去人性化等话语类型标注。
  • 模型在社交媒体数据上训练后无法泛化到特定少数群体场景,反之亦然。
  • 提供非仇恨反例,适合低资源语言仇恨言论研究与合成数据验证。

现有仇恨言论检测系统受限于大规模、细粒度训练数据的缺乏,尤其在英语以外的语言中更为明显。现有语料库多采用简单有毒/无毒标签,而少数涉及特定少数群体的语料又缺少必要的非仇恨反例,难以区分真实仇恨与一般讨论。本文提出 ToxSyn-PT,首个专为葡萄牙语设计的大型多标签仇恨言论数据集,涵盖九类受保护少数群体,包含所有其他公开数据集中缺失的非仇恨反例。数据通过可控的四阶段生成流程构建,并附有话语类型标注,用于捕捉讽刺、去人性化、文化欣赏等修辞策略。实验发现,相较于社交媒体领域数据集,模型在两类任务间存在灾难性互不泛化现象:在社交媒体上训练的模型无法有效迁移到少数群体特定语境,反之亦然。这一结果表明两者是本质不同的任务,且宏观F1等汇总指标可能掩盖模型实际失败,导致评估失真。我们已将 ToxSyn-PT 在 HuggingFace 公开发布,以支持合成数据生成的可复现研究及低资源和中资源语言仇恨言论检测基准进展。

原文摘要 · Abstract (English)

The development of robust hate speech detection systems remains limited by the lack of large-scale, fine-grained training data, especially for languages beyond English. Existing corpora typically rely on simplistic toxic and non-toxic labels, and the few that capture hate directed at specific minority groups lack the positive counterexamples required to distinguish genuine hate from mere discussion. In this work, we introduce ToxSyn-PT, the first Portuguese large-scale corpus explicitly designed for multi-label hate speech detection across nine protected minority groups, including the non-toxic counterexamples absent in all other public datasets. Generated via a controllable four-stage pipeline, ToxSyn contains discourse-type annotations to capture rhetorical strategies of toxic/non-toxic language, such as sarcasm, dehumanization, and cultural appreciation. Our experiments reveal a catastrophic, mutual generalization failure compared to existing datasets from social-media domains: models trained on social media struggle to generalize to minority-specific contexts, and vice-versa. This finding indicates they are distinct tasks and exposes summary metrics like Macro F1 can be unreliable indicators of true model behavior, as they completely mask model failure. We publicly release ToxSyn on HuggingFace to support reproducible research on synthetic data generation and benchmark progress in hate-speech detection for low- and mid-resource languages.

仇恨言论合成数据葡萄牙语多标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。