arXiv:2602.16241cs.CLcs.AI2026-02

测试17个大模型在孟加拉语仇恨言论标注中的表现,发现模型存在偏见且结果不稳定。

Are LLMs Ready to Replace Bangla Annotators?

  • 用统一框架评测17个大模型在零样本下的孟加拉语仇恨言论标注能力
  • 模型规模越大越不稳定,小而对齐的模型反而更一致
  • 适用于低资源语言敏感任务的标注评估,警惕自动化标注风险

大型语言模型(LLMs)正被广泛用于自动化数据标注以加速数据集构建,但在低资源和涉及身份敏感场景中,其作为无偏标注者的可靠性仍不明确。本文系统研究了17个LLM在孟加拉语仇恨言论标注任务中的零样本表现,该任务即使人类标注者也难以达成一致,且标注偏见可能带来严重后果。通过统一评估框架分析发现,模型存在显著标注偏见和判断不稳定性。令人意外的是,模型规模增大并未带来质量提升——更小、更任务对齐的模型往往表现出更强的一致性。这些结果揭示了当前LLMs在低资源语言敏感标注任务中的重要局限,并强调部署前需进行谨慎评估。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used as automated annotators to scale dataset creation, yet their reliability as unbiased annotators--especially for low-resource and identity-sensitive settings--remains poorly understood. In this work, we study the behavior of LLMs as zero-shot annotators for Bangla hate speech, a task where even human agreement is challenging, and annotator bias can have serious downstream consequences. We conduct a systematic benchmark of 17 LLMs using a unified evaluation framework. Our analysis uncovers annotator bias and substantial instability in model judgments. Surprisingly, increased model scale does not guarantee improved annotation quality--smaller, more task-aligned models frequently exhibit more consistent behavior than their larger counterparts. These results highlight important limitations of current LLMs for sensitive annotation tasks in low-resource languages and underscore the need for careful evaluation before deployment.

大模型评估低资源语言标注偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。