arXiv:2608.25605cs.CL2026-08

用指令微调让大模型更好识别仇恨言论,跨领域表现更优。

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

论文配图:From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation
图 1 · 摘自论文原文
  • 通过指令微调通用大模型,统一36个英文仇恨言论数据集进行训练。
  • 在本域任务上达到顶尖水平,跨领域和跨语言泛化能力显著提升。
  • 适合需要强泛化能力的有害内容检测场景,如多语言平台安全审核。

大语言模型(LLMs)在众多通用自然语言任务中表现出色,但在敏感领域如仇恨言论检测中的效果仍不明确。以往研究对比提示型LLMs与主流编码器模型(如BERT变体)发现收益有限,暗示LLMs可能不擅长仇恨言论检测。本文从指令微调视角重新审视该问题,通过整合36个涵盖多种标注方案的英文仇恨言论数据集,对基于Qwen3(Qwen Team, 2025)的通用大模型进行针对性微调,以实现仇恨言论缓解。实验表明,该模型不仅在本域基准上达到当前最优性能,还在跨域和跨语言泛化方面展现出显著优势——这是编码器类专用分类器常遇瓶颈的区域。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate impressive performance across a wide range of general NLP tasks; however, their effectiveness in sensitive domains, such as hate speech detection, remains less clear. Prior studies comparing prompted LLMs with state-of-the-art encoder-based models (e.g., BERT variants (Roy et al., 2023; Dönmez et al., 2024)) have shown only marginal gains, suggesting that LLMs may not excel in hate speech detection or mitigation. In this work, we revisit this question through the lens of instruction tuning. By thoroughly unifying 36 English hate speech datasets spanning multiple labeling schemes, we fine-tune a generalist LLM, based on Qwen3 (Qwen Team, 2025), specifically for hate speech mitigation. Our results demonstrate not only state-of-the-art performance on in-domain benchmarks but also substantial improvements in cross-domain and cross-lingual generalization--areas where encoder-based specialist classifiers often struggle.

大模型仇恨言论指令微调泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。