arXiv:2411.06213cs.CL2024-11

用人类标注的偏见信息提升模型对隐含仇恨言论的识别能力

Incorporating Human Explanations for Robust Hate Speech Detection

  • 引入人类标注的刻板印象、意图和目标群体信息增强上下文理解
  • 设计新任务SIE,使模型能更准确识别隐含的刻板印象意图
  • 实验表明该方法改善内容理解,但隐含意图建模仍存挑战

由于大型Transformer语言模型(LM)存在黑箱特性与复杂性,其泛化能力和鲁棒性在仇恨言论(HS)检测等伦理敏感领域引发担忧。基于富含人类标注的刻板印象、意图和目标群体信息的Social Bias Frames数据集,我们提出三阶段分析:首先发现需建模上下文相关的刻板印象意图以捕捉隐含语义;其次设计新任务Stereotype Intent Entailment(SIE),促使模型在上下文中理解刻板印象的存在;最后通过消融实验与用户研究发现,加入SIE目标可提升模型对内容的理解,但在建模隐含意图方面仍面临挑战。

原文摘要 · Abstract (English)

Given the black-box nature and complexity of large transformer language models (LM), concerns about generalizability and robustness present ethical implications for domains such as hate speech (HS) detection. Using the content rich Social Bias Frames dataset, containing human-annotated stereotypes, intent, and targeted groups, we develop a three stage analysis to evaluate if LMs faithfully assess hate speech. First, we observe the need for modeling contextually grounded stereotype intents to capture implicit semantic meaning. Next, we design a new task, Stereotype Intent Entailment (SIE), which encourages a model to contextually understand stereotype presence. Finally, through ablation tests and user studies, we find a SIE objective improves content understanding, but challenges remain in modeling implicit intent.

仇恨言论检测刻板印象上下文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。