arXiv:2601.20256cs.CL2026-01

构建新基准评估模型对隐性仇恨言论的识别能力

SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility

  • 用论证结构与关联理论生成表面合理但隐含敌意的内容
  • 覆盖7个领域28个群体,共4745条隐性仇恨样本
  • 发现主流模型在隐性仇恨上表现显著下降,适合安全研究者

社交媒体上的仇恨言论分为显性侮辱和威胁(硬仇恨)与表面理性实则引导排斥的隐性仇恨。现有审核系统多针对表层毒害特征优化,对隐性仇恨缺乏鲁棒性,而现有基准未能系统衡量此差距。本文提出 extsc{SoftHateBench},一个生成式基准,通过整合论证主题模型(AMT)与关联理论(RT),将显性仇恨立场重构为看似中立但保持敌意的论述,确保逻辑连贯。该基准涵盖7个社会文化领域、28个目标群体,共4,745条隐性仇恨实例。对编码器类检测器、通用大模型及安全模型的评估显示,系统在从显性到隐性层级时性能普遍下降:能识别明显仇恨的模型,在面对推理驱动的隐蔽表达时往往失效。示例含冒犯性内容,仅用于学术研究。

原文摘要 · Abstract (English)

Online hate on social media ranges from overt slurs and threats (\emph{hard hate speech}) to \emph{soft hate speech}: discourse that appears reasonable on the surface but uses framing and value-based arguments to steer audiences toward blaming or excluding a target group. We hypothesize that current moderation systems, largely optimized for surface toxicity cues, are not robust to this reasoning-driven hostility, yet existing benchmarks do not measure this gap systematically. We introduce \textbf{\textsc{SoftHateBench}}, a generative benchmark that produces soft-hate variants while preserving the underlying hostile standpoint. To generate soft hate, we integrate the \emph{Argumentum Model of Topics} (AMT) and \emph{Relevance Theory} (RT) in a unified framework: AMT provides the backbone argument structure for rewriting an explicit hateful standpoint into a seemingly neutral discussion while preserving the stance, and RT guides generation to keep the AMT chain logically coherent. The benchmark spans \textbf{7} sociocultural domains and \textbf{28} target groups, comprising \textbf{4,745} soft-hate instances. Evaluations across encoder-based detectors, general-purpose LLMs, and safety models show a consistent drop from hard to soft tiers: systems that detect explicit hostility often fail when the same stance is conveyed through subtle, reasoning-based language. \textcolor{red}{\textbf{Disclaimer.} Contains offensive examples used solely for research.}

仇恨言论隐性攻击评估基准大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。