用大模型辅助识别仇恨言论中的可核查内容,提升检测效率与准确性。
When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech
- 构建人类标注+大模型协作的标注框架,降低人工成本。
- 在WSF-ARG+数据集上,大模型检测准确率提升0.213(宏F1)。
- 适合从事有害内容检测、信息真实性验证的研究者使用。
网络仇恨言论常以看似事实的形式出现,尤其在有组织的骚扰和极端主义宣传中。若不同时处理仇恨言论与虚假信息,将加剧偏见、固化刻板印象,并使旁观者遭受心理伤害,污染公共讨论。此类内容需审核人员同时判断其危害性与真实性,工作量更大。为此,我们发布首个融合仇恨言论与可核查性信息的数据集WSF-ARG+,并提出一种新型大模型协同标注框架。通过12种不同规模与架构的开源大模型测试,经人工评估验证,该框架显著减少人力投入且不降低标注质量。实验表明,含有可核查声明的仇恨言论具有更高攻击性,引入可核查标签后,大模型在仇恨言论检测上的宏F1值平均提升0.154,最大提升达0.213。
原文摘要 · Abstract (English)
Hateful content online is often expressed using fact-like, not necessarily correct information, especially in coordinated online harassment campaigns and extremist propaganda. Failing to jointly address hate speech (HS) and misinformation can deepen prejudice, reinforce harmful stereotypes, and expose bystanders to psychological distress, while polluting public debate. Moreover, these messages require more effort from content moderators because they must assess both harmfulness and veracity, i.e., fact-check them. To address this challenge, we release WSF-ARG+, the first dataset which combines hate speech with check-worthiness information. We also introduce a novel LLM-in-the-loop framework to facilitate the annotation of check-worthy claims. We run our framework, testing it with 12 open-weight LLMs of different sizes and architectures. We validate it through extensive human evaluation, and show that our LLM-in-the-loop framework reduces human effort without compromising the annotation quality of the data. Finally, we show that HS messages with check-worthy claims show significantly higher harassment and hate, and that incorporating check-worthiness labels improves LLM-based HS detection up to 0.213 macro-F1 and to 0.154 macro-F1 on average for large models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。