arXiv:2502.06207cs.CLcs.AI2025-02ACL被引 21

LLM在争议性内容判断中容易过度自信,但用争议样本训练可提升准确率。

Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement

  • 用争议样本训练,让LLM学会处理模糊判断。
  • 低一致性的内容上,模型准确率下降且信心过高。
  • 适合内容审核系统优化与模型可靠性研究者。

大型语言模型(LLMs)已成为检测不当语言的关键工具,但其在处理标注分歧方面的表现仍不明确。由于主观理解差异产生的分歧样本具有高度模糊性,构成独特挑战。本研究系统评估了多种LLM在不同标注一致性水平下的表现,分析二分类准确率,考察模型置信度与人类分歧之间的关系,并探究分歧样本对少样本学习和指令微调中决策的影响。结果表明,LLM在低一致性的样本上表现不佳,常对模糊情况表现出过度自信。然而,将分歧样本纳入训练能同时提高检测准确率和模型与人类判断的一致性。这些发现为实际内容审核中基于LLM的不当语言检测提供了优化基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become essential for offensive language detection, yet their ability to handle annotation disagreement remains underexplored. Disagreement samples, which arise from subjective interpretations, pose a unique challenge due to their ambiguous nature. Understanding how LLMs process these cases, particularly their confidence levels, can offer insight into their alignment with human annotators. This study systematically evaluates the performance of multiple LLMs in detecting offensive language at varying levels of annotation agreement. We analyze binary classification accuracy, examine the relationship between model confidence and human disagreement, and explore how disagreement samples influence model decision-making during few-shot learning and instruction fine-tuning. Our findings reveal that LLMs struggle with low-agreement samples, often exhibiting overconfidence in these ambiguous cases. However, utilizing disagreement samples in training improves both detection accuracy and model alignment with human judgment. These insights provide a foundation for enhancing LLM-based offensive language detection in real-world moderation tasks.

大模型评估内容安全标注分歧模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。