跨文化分析仇恨言论定义差异,揭示其语义复杂性与检测敏感性。
Untangling Hate Speech Definitions: A Semantic Componential Analysis Across Cultures and Domains
- 提出语义成分分析框架,拆解跨文化仇恨言论定义。
- 构建首个包含493条定义的多文化数据集,覆盖五大领域。
- 发现大模型对定义复杂度敏感,提示需考虑文化适配性。
仇恨言论高度依赖文化背景,导致个体理解差异显著。为此,我们提出语义成分分析(SCA)框架,开展跨文化、跨领域的仇恨言论定义分析。我们构建了首个涵盖493条定义的数据集,源自100多个文化背景,覆盖在线词典、学术研究、维基百科、法律文本及在线平台五大领域。通过将定义分解为语义成分,分析显示定义间存在显著差异,但多个领域在引用时未充分考虑目标文化背景。我们利用三个主流开源大语言模型(LLMs)进行零样本实验,探究不同定义对仇恨言论检测的影响。结果表明,大模型对定义敏感:提示中使用定义的复杂度直接影响检测响应。这凸显了在构建检测系统时需重视文化语境与定义一致性。
原文摘要 · Abstract (English)
Hate speech relies heavily on cultural influences, leading to varying individual interpretations. For that reason, we propose a Semantic Componential Analysis (SCA) framework for a cross-cultural and cross-domain analysis of hate speech definitions. We create the first dataset of hate speech definitions encompassing 493 definitions from more than 100 cultures, drawn from five key domains: online dictionaries, academic research, Wikipedia, legal texts, and online platforms. By decomposing these definitions into semantic components, our analysis reveals significant variation across definitions, yet many domains borrow definitions from one another without taking into account the target culture. We conduct zero-shot model experiments using our proposed dataset, employing three popular open-sourced LLMs to understand the impact of different definitions on hate speech detection. Our findings indicate that LLMs are sensitive to definitions: responses for hate speech detection change according to the complexity of definitions used in the prompt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。