arXiv:2506.12744cs.CLcs.CY2025-06被引 8

大模型在社交媒体仇恨言论检测中表现超越传统模型。

Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?

  • 用大语言模型替代专用BERT模型进行检测
  • 在少数据下仍优于微调过的BERT模型
  • 适合多语言混用场景下的研究者参考

当前社交媒体中的仇恨言论检测面临语言多样性与非正式表达的挑战,尤其在涉及印地语-英语混用、音译及文化特有表达时更为突出。尽管基于BERT的微调模型已成为标准方法,本文认为最新的大语言模型(LLMs)不仅性能更优,更重新定义了该任务的范式。为此,我们构建了IndoHateMix——一个高质量、多样化的印度语境下印地语-英语混用与音译数据集,为复杂多语言场景提供真实基准。大量实验表明,前沿大模型(如LLaMA-3.1)即使在远少于训练数据的情况下,也持续优于专门微调的BERT模型。其卓越的泛化与适应能力,展现出在多元环境治理网络仇恨言论的变革性潜力。这引发思考:未来研究应聚焦专用模型开发,还是投入更丰富多样的数据集以进一步提升大模型效能?

原文摘要 · Abstract (English)

Hate speech detection across contemporary social media presents unique challenges due to linguistic diversity and the informal nature of online discourse. These challenges are further amplified in settings involving code-mixing, transliteration, and culturally nuanced expressions. While fine-tuned transformer models, such as BERT, have become standard for this task, we argue that recent large language models (LLMs) not only surpass them but also redefine the landscape of hate speech detection more broadly. To support this claim, we introduce IndoHateMix, a diverse, high-quality dataset capturing Hindi-English code-mixing and transliteration in the Indian context, providing a realistic benchmark to evaluate model robustness in complex multilingual scenarios where existing NLP methods often struggle. Our extensive experiments show that cutting-edge LLMs (such as LLaMA-3.1) consistently outperform task-specific BERT-based models, even when fine-tuned on significantly less data. With their superior generalization and adaptability, LLMs offer a transformative approach to mitigating online hate in diverse environments. This raises the question of whether future works should prioritize developing specialized models or focus on curating richer and more varied datasets to further enhance the effectiveness of LLMs.

仇恨言论检测大语言模型多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。