arXiv:2505.10260cs.CLcs.AI2025-05

测试大模型识别社交媒体中人权侵犯内容的能力,比较其准确率与语言适应性。

Comparing LLM Text Annotation Skills: A Study on Human Rights Violations in Social Media Data

  • 用零样本和少样本提示测试5个大模型对俄语/乌克兰语社交文本的标注能力。
  • 在1000个样本上对比模型与人工双标注结果,评估分类准确率差异。
  • 揭示各模型在跨语言场景下的错误模式,适合关注敏感内容检测的研究者。

在自然语言处理日益复杂的背景下,大语言模型(LLMs)展现出在需要细致文本理解与上下文推理任务中的巨大潜力。本研究考察了GPT-3.5、GPT-4、LLAMA3、Mistral 7B和Claude-2等前沿模型在零样本与少样本条件下,对包含俄语和乌克兰语社交媒体帖子的复杂数据集进行人权侵犯内容二分类标注的能力。通过将模型标注结果与1000个样本的人工双标注黄金标准进行对比,评估其性能。研究还分析了不同提示语境(英语与俄语)下的表现差异,并探讨各模型特有的错误模式与分歧特征,揭示其优势、局限及跨语言适应性。该工作有助于理解大模型在多语言敏感领域任务中的可靠性与适用性,尤其对主观性强、依赖上下文的判断机制提供实证洞察,为实际部署提供参考。

原文摘要 · Abstract (English)

In the era of increasingly sophisticated natural language processing (NLP) systems, large language models (LLMs) have demonstrated remarkable potential for diverse applications, including tasks requiring nuanced textual understanding and contextual reasoning. This study investigates the capabilities of multiple state-of-the-art LLMs - GPT-3.5, GPT-4, LLAMA3, Mistral 7B, and Claude-2 - for zero-shot and few-shot annotation of a complex textual dataset comprising social media posts in Russian and Ukrainian. Specifically, the focus is on the binary classification task of identifying references to human rights violations within the dataset. To evaluate the effectiveness of these models, their annotations are compared against a gold standard set of human double-annotated labels across 1000 samples. The analysis includes assessing annotation performance under different prompting conditions, with prompts provided in both English and Russian. Additionally, the study explores the unique patterns of errors and disagreements exhibited by each model, offering insights into their strengths, limitations, and cross-linguistic adaptability. By juxtaposing LLM outputs with human annotations, this research contributes to understanding the reliability and applicability of LLMs for sensitive, domain-specific tasks in multilingual contexts. It also sheds light on how language models handle inherently subjective and context-dependent judgments, a critical consideration for their deployment in real-world scenarios.

大模型文本标注人权监测多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。