arXiv:2604.14970cs.CL2026-04ACL

用混合方法解释仇恨言论,让平台更透明地说明为何内容被标记

Explain the Flag: Contextualizing Hate Speech Beyond Censorship

论文配图:Explain the Flag: Contextualizing Hate Speech Beyond Censorship
图 1 · 摘自论文原文
  • 结合大模型与自建词汇库,双路径识别歧视性用语和群体攻击内容
  • 在英法希三语中实现高准确率检测,生成可理解的标注理由
  • 适合需要透明审核的社交平台、内容治理团队使用

仇恨、贬损和冒犯性言论仍是在线平台和公共讨论中的长期挑战。尽管自动化检测系统广泛应用,但多数聚焦于删帖或屏蔽,引发透明度与表达自由的担忧,也限制了对内容危害原因的解释。为此,解释性方法应运而生,旨在使仇恨言论检测更具透明度、可问责性和信息量。本文提出一种混合方法,结合大型语言模型(LLMs)与三个新创建并精心筛选的词汇库,用于英语、法语和希腊语的仇恨言论检测与解释。系统通过两条互补路径捕捉两类内容:一是利用词汇库识别并消歧具有身份指向性的贬损表达;二是借助大模型作为上下文感知的群体攻击评估器。最终输出融合为有依据的解释,说明内容为何被标记。人工评估表明,该混合方法在准确性与解释质量上均优于仅使用大模型的基线。

原文摘要 · Abstract (English)

Hate, derogatory, and offensive speech remains a persistent challenge in online platforms and public discourse. While automated detection systems are widely used, most focus on censorship or removal, raising concerns for transparency and freedom of expression, and limiting opportunities to explain why content is harmful. To address these issues, explanatory approaches have emerged as a promising solution, aiming to make hate speech detection more transparent, accountable, and informative. In this paper, we present a hybrid approach that combines Large Language Models (LLMs) with three newly created and curated vocabularies to detect and explain hate speech in English, French, and Greek. Our system captures both inherently derogatory expressions tied to identity characteristics and direct group-targeted content through two complementary pipelines: one that detects and disambiguates problematic terms using the curated vocabularies, and one that leverages LLMs as context-aware evaluators of group-targeting content. The outputs are fused into grounded explanations that clarify why content is flagged. Human evaluation shows that our hybrid approach is accurate, with high-quality explanations, outperforming LLM-only baselines.

仇恨言论可解释性多语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。