用思维链蒸馏大模型,让小模型既快又会解释仇恨言论。
Towards Efficient and Explainable Hate Speech Detection via Model Distillation
- 用思维链提取大模型解释,蒸馏成小型可解释模型
- 小模型分类准确率超过原大模型,解释质量相当
- 适合需要实时、透明检测的平台和监管场景
自动检测仇恨与辱骂性语言对遏制其在线传播至关重要。同时,识别并解释仇恨言论有助于让人们认识其负面影响。然而,当前大多数检测模型是黑箱,缺乏可解释性。大型语言模型(LLMs)在仇恨言论检测中表现有效,并能提升可解释性,但运行成本高。本文提出一种基于思维链(Chain-of-Thought)的模型蒸馏方法,将大模型中的解释能力迁移至小型模型。实验表明,蒸馏后的小模型在分类性能上优于原大模型,同时保持同等水平的解释质量。这一双重优势使仇恨言论检测更高效、可理解且可操作,适用于实际部署场景。
原文摘要 · Abstract (English)
Automatic detection of hate and abusive language is essential to combat its online spread. Moreover, recognising and explaining hate speech serves to educate people about its negative effects. However, most current detection models operate as black boxes, lacking interpretability and explainability. In this context, Large Language Models (LLMs) have proven effective for hate speech detection and to promote interpretability. Nevertheless, they are computationally costly to run. In this work, we propose distilling big language models by using Chain-of-Thought to extract explanations that support the hate speech classification task. Having small language models for these tasks will contribute to their use in operational settings. In this paper, we demonstrate that distilled models deliver explanations of the same quality as larger models while surpassing them in classification performance. This dual capability, classifying and explaining, advances hate speech detection making it more affordable, understandable and actionable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。