arXiv:2506.12148cs.CL2025-06ACL

静态基准无法反映仇恨言论随时间演化的真实风险,需动态评估。

Hatevolution: What Static Benchmarks Don't Tell Us

  • 构建两个随时间演化的仇恨言论实验,测试模型鲁棒性。
  • 20个语言模型在静态与动态评估中表现差异显著。
  • 呼吁建立时序敏感的评测基准,提升模型安全性评估可靠性。

语言随时间演变,仇恨言论领域尤其受社会动态和文化变迁影响,快速演化。尽管已有研究探讨语言演变对模型训练的影响并提出若干解决方案,但其对模型评测的影响仍被忽视。然而,仇恨言论评测对确保模型安全至关重要。本文通过两个演化型仇恨言论实验,实证评估了20个语言模型的鲁棒性,揭示了静态评测与时间敏感评测之间存在明显的时间错位。研究结果表明,亟需建立时序敏感的语言评测基准,以更准确、可靠地评估语言模型在仇恨言论领域的表现。

原文摘要 · Abstract (English)

Language changes over time, including in the hate speech domain, which evolves quickly following social dynamics and cultural shifts. While NLP research has investigated the impact of language evolution on model training and has proposed several solutions for it, its impact on model benchmarking remains under-explored. Yet, hate speech benchmarks play a crucial role to ensure model safety. In this paper, we empirically evaluate the robustness of 20 language models across two evolving hate speech experiments, and we show the temporal misalignment between static and time-sensitive evaluations. Our findings call for time-sensitive linguistic benchmarks in order to correctly and reliably evaluate language models in the hate speech domain.

仇恨言论动态评测语言演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。