arXiv:2508.20750cs.CL2025-08中稿 · the DHOW Workshop …被引 3

微调通用大模型嵌入,显著提升隐性仇恨言论跨数据集检测效果。

Specializing General-purpose LLM Embeddings for Implicit Hate Speech Detection across Datasets

  • 仅通过微调通用大模型嵌入,无需额外知识
  • 跨数据集检测F1-macro最高提升20.35个百分点
  • 适合需高效迁移的隐性仇恨言论识别场景

隐性仇恨言论(IHS)通过微妙线索、讽刺或编码术语表达偏见或敌意,不包含明确侮辱性词汇,因而难以检测。本文表明,仅通过微调基于大语言模型的通用嵌入模型(如Stella、Jasper、NV-Embed和E5),即可实现最佳性能。在多个IHS数据集上的实验显示,同数据集评估中F1-macro最高提升1.10个百分点,跨数据集评估中最高提升20.35个百分点。

原文摘要 · Abstract (English)

Implicit hate speech (IHS) is indirect language that conveys prejudice or hatred through subtle cues, sarcasm or coded terminology. IHS is challenging to detect as it does not include explicit derogatory or inflammatory words. To address this challenge, task-specific pipelines can be complemented with external knowledge or additional information such as context, emotions and sentiment data. In this paper, we show that, by solely fine-tuning recent general-purpose embedding models based on large language models (LLMs), such as Stella, Jasper, NV-Embed and E5, we achieve state-of-the-art performance. Experiments on multiple IHS datasets show up to 1.10 percentage points improvements for in-dataset, and up to 20.35 percentage points improvements in cross-dataset evaluation, in terms of F1-macro score.

仇恨言论嵌入微调大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。