arXiv:2601.03194cs.CL2026-01AAAI被引 6

针对印地语等小语种仇恨言论,提出可解释检测新框架并构建多语言标注数据集。

X-MuTeST: A Multilingual Benchmark for Explainable Hate Speech Detection and A Novel LLM-consulted Explanation Framework

  • 融合大模型语义推理与注意力机制,提升检测可解释性。
  • 在3种语言共1.6万条数据上验证,解释力指标提升显著。
  • 适合关注多语言仇恨言论与AI可解释性的研究者使用。

社交媒体上的仇恨言论检测面临准确率与可解释性双重挑战,尤其在资源匮乏的印地语等印地语系语言中更为突出。本文提出X-MuTeST(可解释多语言仇恨言论检测)框架,结合大语言模型(LLM)的高层语义推理与传统注意力增强技术,以提升模型性能与可解释性。研究扩展至印地语、泰卢固语及英语,并提供每词级别的人工标注理由以支持分类标签。该方法通过比较原始文本与单字、双字、三字组预测概率差异生成解释,最终解释为LLM解释与X-MuTeST解释的并集。实验表明,利用人工理由训练可同时提升分类性能与可解释性;进一步结合人工理由优化模型注意力,效果更优。通过Token-F1、IOU-F1等可理解性指标及完备性、充分性等忠实度指标评估解释质量。本研究涵盖6,004条印地语、4,492条泰卢固语与6,334条英语样本的粒度级理由标注。代码与数据已开源。

原文摘要 · Abstract (English)

Hate speech detection on social media faces challenges in both accuracy and explainability, especially for underexplored Indic languages. We propose a novel explainability-guided training framework, X-MuTeST (eXplainable Multilingual haTe Speech deTection), for hate speech detection that combines high-level semantic reasoning from large language models (LLMs) with traditional attention-enhancing techniques. We extend this research to Hindi and Telugu alongside English by providing benchmark human-annotated rationales for each word to justify the assigned class label. The X-MuTeST explainability method computes the difference between the prediction probabilities of the original text and those of unigrams, bigrams, and trigrams. Final explanations are computed as the union between LLM explanations and X-MuTeST explanations. We show that leveraging human rationales during training enhances both classification performance and explainability. Moreover, combining human rationales with our explainability method to refine the model attention yields further improvements. We evaluate explainability using Plausibility metrics such as Token-F1 and IOU-F1 and Faithfulness metrics such as Comprehensiveness and Sufficiency. By focusing on under-resourced languages, our work advances hate speech detection across diverse linguistic contexts. Our dataset includes token-level rationale annotations for 6,004 Hindi, 4,492 Telugu, and 6,334 English samples. Data and code are available on https://github.com/ziarehman30/X-MuTeST

仇恨言论多语言可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。