arXiv:2507.10468cs.CLcs.LG2025-07

对比经典模型与大语言模型,验证其在真实网络仇恨言论检测中的表现

From BERT to Qwen: Hate Detection across architectures

  • 在真实在线文本上对比编码器与自回归大模型的检测能力
  • 发现大模型在复杂语境下识别率提升,但误判率也更高
  • 适合关注生成式模型实际应用效果的研究者参考

在线平台在遏制仇恨言论时常面临过度审查合法讨论的问题。早期双向Transformer编码器取得了显著进展,而超大规模自回归大语言模型则有望实现更深层的上下文感知。然而,这种规模优势是否真正提升真实文本中的仇恨言论检测效果,尚未得到验证。本研究通过在精心构建的在线互动语料库上对两类模型——经典编码器与新一代大语言模型——进行基准测试,评估其在‘仇恨’或‘非仇恨’分类任务上的表现。

原文摘要 · Abstract (English)

Online platforms struggle to curb hate speech without over-censoring legitimate discourse. Early bidirectional transformer encoders made big strides, but the arrival of ultra-large autoregressive LLMs promises deeper context-awareness. Whether this extra scale actually improves practical hate-speech detection on real-world text remains unverified. Our study puts this question to the test by benchmarking both model families, classic encoders and next-generation LLMs, on curated corpora of online interactions for hate-speech detection (Hate or No Hate).

仇恨言论检测大语言模型文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。