研究种族身份如何影响仇恨言论判断,揭示了群体差异的语义模式。
Semantic Gradients Interactions in SSD: A Case Study in Racial Identity and Hate Speech
- 扩展语义差异法,建模语义随群体等条件变化的交互效应。
- 在仇恨言论数据集上发现种族身份显著调节判断,但差异较小。
- 方法可解释、可检验,适合社会认知与偏见研究者使用。
我们提出交互式语义差异法(Interaction SSD),扩展了监督语义差异法,用于建模语义意义如何随群体、特质或条件的变化而变化,使这种变化可检验且可解释。该方法估计主语义梯度、交互梯度和条件梯度,均可用标准SSD工具解析。以加州大学伯克利分校仇恨言论标注语料库为例,检验标注者种族身份是否调节对针对有色人种评论的仇恨言论判断。结果显示存在显著调节效应:共同梯度对比去人性化敌意与反言辞,而交互梯度揭示了群体相关的小幅差异,即不同语义线索预测仇恨言论评分的程度各异。交互SSD使受调制的意义-结果关系具备统计可检验性与可解释性。
原文摘要 · Abstract (English)
We introduce interaction SSD, an extension of Supervised Semantic Differential that models how semantic meaning varies across moderators such as groups, traits, or conditions making this variation testable and interpretable. The method estimates a main semantic gradient, an interaction gradient, and conditional gradients, all interpretable through standard SSD tools. We illustrate it on the UC Berkeley Measuring Hate Speech corpus, testing whether annotator racial identity moderates hate-speech judgments of comments targeting people of color. The interaction model detects a significant moderation effect: the shared gradient contrasts dehumanizing hostility with counter-speech, while the interaction gradient reveals smaller group-linked differences in which semantic cues predict hate-speech ratings. Interaction SSD makes moderated meaning-outcome relationships statistically testable and interpretable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。