LLMs虽不能准确标注仇恨言论,但能可靠评估模型优劣排序。
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
- 用更合理的主观性框架衡量LLM与人类的差异
- LLM标注虽不一致,但能保持模型性能排名一致
- 适合需要快速评估模型表现的场景
仇恨言论在线广泛传播,危害个人与社群,自动检测对大规模内容治理至关重要,但识别仍具挑战。部分难点源于主观性:不同人对同一内容判断不一。传统标注一致性指标(如Cohen's κ)将分歧视为错误,未体现其意义。尽管大语言模型(LLMs)具备规模化标注潜力,但先前研究显示其在主观任务中难以替代人工判断。本文采用主观性感知框架——跨标注者可靠性(xRR),重新评估LLM可靠性,发现即使在更公平视角下,LLM仍与人类存在差异。然而这一局限带来新机遇:我们发现LLM生成的标注可稳定反映不同分类模型的性能趋势,与人工评价高度相关。通过检验模型在人工标注下的相对性能排序是否在LLM标注下保持一致,结果表明:尽管实例级判断存在偏差,但模型排序模式高度相似,暗示其作为代理评估者的潜力。因此,虽无法取代人类标注,但在主观自然语言处理任务中可作为可扩展的评估代理。
原文摘要 · Abstract (English)
Hate speech spreads widely online, harming individuals and communities, making automatic detection essential for large-scale moderation, yet detecting it remains difficult. Part of the challenge lies in subjectivity: what one person flags as hate speech, another may see as benign. Traditional annotation agreement metrics, such as Cohen's $κ$, oversimplify this disagreement, treating it as an error rather than meaningful diversity. Meanwhile, Large Language Models (LLMs) promise scalable annotation, but prior studies demonstrate that they cannot fully replace human judgement, especially in subjective tasks. In this work, we reexamine LLM reliability using a subjectivity-aware framework, cross-Rater Reliability (xRR), revealing that even under fairer lens, LLMs still diverge from humans. Yet this limitation opens an opportunity: we find that LLM-generated annotations can reliably reflect performance trends across classification models, correlating with human evaluations. We test this by examining whether LLM-generated annotations preserve the relative ordering of model performance derived from human evaluation (i.e. whether models ranked as more reliable by human annotators preserve the same order when evaluated with LLM-generated labels). Our results show that, although LLMs differ from humans at the instance level, they reproduce similar ranking and classification patterns, suggesting their potential as proxy evaluators. While not a substitute for human annotators, they might serve as a scalable proxy for evaluation in subjective NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。