arXiv:2506.18576cs.CLcs.CY2025-06被引 9

构建仇恨言论定义的模块化分类体系,揭示不同定义对大模型零样本识别效果的影响。

A Modular Taxonomy for Hate Speech Definitions and Its Impact on Zero-Shot LLM Classification Performance

  • 提出14个概念要素构成的仇恨言论定义分类体系
  • 不同定义下大模型性能差异显著,但效果不一致
  • 适用于研究仇恨言论定义与模型性能关系的学者

检测有害内容是自然语言处理用于社会公益的关键任务,其中仇恨言论尤为危险。然而,何为仇恨言论?如何定义?不同定义的提示如何影响模型表现?本文贡献有二:理论层面,通过梳理文献中的现有定义,构建由14个概念要素组成的分类体系,涵盖目标对象(个人或群体)及潜在后果等维度;实验层面,基于该体系对三类大模型在三个不同数据集(合成、人机协同、真实场景)上进行系统性零样本评估。结果表明,采用不同具体程度的定义会影响模型表现,但这种影响在不同架构间并不一致。

原文摘要 · Abstract (English)

Detecting harmful content is a crucial task in the landscape of NLP applications for Social Good, with hate speech being one of its most dangerous forms. But what do we mean by hate speech, how can we define it, and how does prompting different definitions of hate speech affect model performance? The contribution of this work is twofold. At the theoretical level, we address the ambiguity surrounding hate speech by collecting and analyzing existing definitions from the literature. We organize these definitions into a taxonomy of 14 Conceptual Elements-building blocks that capture different aspects of hate speech definitions, such as references to the target of hate (individual or groups) or of the potential consequences of it. At the experimental level, we employ the collection of definitions in a systematic zero-shot evaluation of three LLMs, on three hate speech datasets representing different types of data (synthetic, human-in-the-loop, and real-world). We find that choosing different definitions, i.e., definitions with a different degree of specificity in terms of encoded elements, impacts model performance, but this effect is not consistent across all architectures.

仇恨言论大模型评估定义体系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。