arXiv:2411.17637cs.CLcs.LG2024-11被引 23

LLM当标注员在低资源语言上效果差,还不如微调的BERT模型。

On Limitations of LLM as Annotator for Low Resource Languages

  • 用大模型生成标注数据,但对马拉地语效果不佳。
  • 最先进大模型比微调BERT低10%以上准确率。
  • 适合关注低资源语言NLP落地的研究者看。

低资源语言因缺乏足够的语言数据、资源和工具,在监督学习、标注和分类等任务中面临重大挑战,阻碍了情感分析、仇恨言论检测等关键NLP任务的发展。为弥合这一差距,大型语言模型(LLMs)被视为潜在的标注工具,可生成相关数据集与资源。本文聚焦马拉地语这一低资源语言,评估闭源与开源LLMs作为标注器的表现,并与微调的BERT模型对比。测试模型包括GPT-4o、Gemini 1.0 Pro、Gemma 2(2B和9B)、Llama 3.1(8B和405B),涵盖情感分析、新闻分类和仇恨言论检测任务。结果表明,尽管大模型在高资源语言如英语中表现优异,但在马拉地语上仍明显不足。即使最先进的GPT-4o和Llama 3.1 405B也低于微调的BERT基线,准确率分别落后10.2%和14.1%,凸显了大模型作为低资源语言标注工具的局限性。

原文摘要 · Abstract (English)

Low-resource languages face significant challenges due to the lack of sufficient linguistic data, resources, and tools for tasks such as supervised learning, annotation, and classification. This shortage hinders the development of accurate models and datasets, making it difficult to perform critical NLP tasks like sentiment analysis or hate speech detection. To bridge this gap, Large Language Models (LLMs) present an opportunity for potential annotators, capable of generating datasets and resources for these underrepresented languages. In this paper, we focus on Marathi, a low-resource language, and evaluate the performance of both closed-source and open-source LLMs as annotators, while also comparing these results with fine-tuned BERT models. We assess models such as GPT-4o and Gemini 1.0 Pro, Gemma 2 (2B and 9B), and Llama 3.1 (8B and 405B) on classification tasks including sentiment analysis, news classification, and hate speech detection. Our findings reveal that while LLMs excel in annotation tasks for high-resource languages like English, they still fall short when applied to Marathi. Even advanced models like GPT-4o and Llama 3.1 405B underperform compared to fine-tuned BERT-based baselines, with GPT-4o and Llama 3.1 405B trailing fine-tuned BERT by accuracy margins of 10.2% and 14.1%, respectively. This highlights the limitations of LLMs as annotators for low-resource languages.

低资源语言大模型标注马拉地语模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。