arXiv:2508.08274cs.CLcs.AI2025-08中稿 · publication in Inf…被引 5

用形容词做可解释的中间表示,提升仇恨与反仇恨内容识别效果

Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition

  • 用大模型将文本转为形容词表示,再由轻量分类器判断
  • 跨五数据集平均宏F1达0.69,四组优于现有方法
  • 结果可解释性强,适配其他NLP任务

社交媒体上仇恨言论的激增对社会造成前所未有的影响,自动化检测方法因而至关重要。与以往黑箱模型不同,本文提出一种透明的自动仇恨与反仇恨言论识别方法——“话语概念瓶颈模型”(SCBM),以形容词作为人类可理解的瓶颈概念。SCBM利用大语言模型(LLMs)将输入文本映射为抽象的形容词表示,再送入轻量分类器完成下游任务。在涵盖多种语言和平台(如Twitter、Reddit、YouTube)的五个基准数据集上,SCBM平均宏F1得分为0.69,优于文献中最新结果的四组数据。除高准确率外,该方法还具备较强的局部与全局可解释性。进一步融合形容词概念表示与Transformer嵌入,各数据集平均性能提升1.8%,表明该表示能捕捉互补信息。结果表明,基于形容词的概念表示可作为仇恨与反仇恨言论识别的紧凑、可解释且有效的编码方式。通过调整形容词,本方法亦可扩展至其他NLP任务。

原文摘要 · Abstract (English)

The rapid increase in hate speech on social media has exposed an unprecedented impact on society, making automated methods for detecting such content important. Unlike prior black-box models, we propose a novel transparent method for automated hate and counter speech recognition, i.e., "Speech Concept Bottleneck Model" (SCBM), using adjectives as human-interpretable bottleneck concepts. SCBM leverages large language models (LLMs) to map input texts to an abstract adjective-based representation, which is then sent to a light-weight classifier for downstream tasks. Across five benchmark datasets spanning multiple languages and platforms (e.g., Twitter, Reddit, YouTube), SCBM achieves an average macro-F1 score of 0.69 which outperforms the most recently reported results from the literature on four out of five datasets. Aside from high recognition accuracy, SCBM provides a high level of both local and global interpretability. Furthermore, fusing our adjective-based concept representation with transformer embeddings, leads to a 1.8% performance increase on average across all datasets, showing that the proposed representation captures complementary information. Our results demonstrate that adjective-based concept representations can serve as compact, interpretable, and effective encodings for hate and counter speech recognition. With adapted adjectives, our method can also be applied to other NLP tasks.

可解释性仇恨言论概念瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。