构建多维可解释数据集,揭示虚假言论如何隐性煽动仇恨
HateMirage: An Explainable Multi-Dimensional Dataset for Decoding Faux Hate and Subtle Online Abuse
- 基于被辟谣信息生成4530条假仇恨评论,标注目标、意图与影响三维度
- 模型解释质量更依赖预训练多样性而非规模,小模型也可表现优异
- 适合研究虚假信息与仇恨传播关系的AI安全与可解释性学者
隐蔽且间接的仇恨言论仍是在线安全研究中的未充分探索问题,尤其当有害意图隐藏在误导性或操纵性叙事中时。现有仇恨言论数据集主要捕捉明显攻击性内容,低估了错误信息如何煽动或合理化仇恨的复杂方式。为此,我们提出HateMirage,一个专为虚假仇恨(Faux Hate)设计的新数据集,旨在推动对虚假叙事中仇恨生成机制的推理与可解释性研究。该数据集通过追踪事实核查来源中的广泛辟谣信息及其相关YouTube讨论,收集了4,530条用户评论。每条评论均从三个可解释维度进行标注:目标(受影响对象)、意图(评论背后动机或目的)、影响(潜在社会后果)。与HateXplain、HARE等仅提供词级或单维度解释的数据集不同,HateMirage引入多维度解释框架,捕捉错误信息、伤害行为与社会后果之间的互动关系。我们在HateMirage上对多个开源语言模型进行基准测试,使用ROUGE-L F1和Sentence-BERT相似度评估解释连贯性。结果表明,解释质量可能更多取决于预训练数据的多样性与推理导向数据,而非模型规模本身。通过将错误信息推理与伤害归因相结合,HateMirage为可解释仇恨检测与负责任的AI研究设立了新基准。
原文摘要 · Abstract (English)
Subtle and indirect hate speech remains an underexplored challenge in online safety research, particularly when harmful intent is embedded within misleading or manipulative narratives. Existing hate speech datasets primarily capture overt toxicity, underrepresenting the nuanced ways misinformation can incite or normalize hate. To address this gap, we present HateMirage, a novel dataset of Faux Hate comments designed to advance reasoning and explainability research on hate emerging from fake or distorted narratives. The dataset was constructed by identifying widely debunked misinformation claims from fact-checking sources and tracing related YouTube discussions, resulting in 4,530 user comments. Each comment is annotated along three interpretable dimensions: Target (who is affected), Intent (the underlying motivation or goal behind the comment), and Implication (its potential social impact). Unlike prior explainability datasets such as HateXplain and HARE, which offer token-level or single-dimensional reasoning, HateMirage introduces a multi-dimensional explanation framework that captures the interplay between misinformation, harm, and social consequence. We benchmark multiple open-source language models on HateMirage using ROUGE-L F1 and Sentence-BERT similarity to assess explanation coherence. Results suggest that explanation quality may depend more on pretraining diversity and reasoning-oriented data rather than on model scale alone. By coupling misinformation reasoning with harm attribution, HateMirage establishes a new benchmark for interpretable hate detection and responsible AI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。