arXiv:2510.12807cs.CLcs.AI2025-10被引 1

评测开源大模型在波斯语零样本与少样本学习表现

Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning

  • 用零样本和少样本框架测试多个开源模型在波斯语任务上的表现
  • Gemma 2 在多数任务中领先,尤其在复杂推理上表现突出
  • 模型在命名实体识别等细粒度任务上普遍表现不佳

大型语言模型在多种语言上展现出卓越能力,但在低资源语言如波斯语中的效果仍需深入研究。本文对多个开源大模型在波斯语自然语言处理任务上进行了全面基准测试,涵盖情感分析、命名实体识别、阅读理解与问答等任务,使用ParsiNLU和ArmanEmo等标准波斯语数据集。实验采用零样本与少样本学习范式,以准确率、F1分数、BLEU和ROUGE等指标评估性能。结果显示,Gemma 2在几乎所有任务中均优于其他模型,尤其在复杂推理任务中表现显著;但多数模型在命名实体识别等词级别理解任务上表现较差,凸显波斯语处理的挑战。本研究为多语言大模型在波斯语中的应用提供了重要参考,也为未来模型开发建立了基准。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities across numerous languages; however, their effectiveness in low-resource languages like Persian requires thorough investigation. This paper presents a comprehensive benchmark of several open-source LLMs for Persian Natural Language Processing (NLP) tasks, utilizing both zero-shot and few-shot learning paradigms. We evaluate models across a range of tasks including sentiment analysis, named entity recognition, reading comprehension, and question answering, using established Persian datasets such as ParsiNLU and ArmanEmo. Our methodology encompasses rigorous experimental setups for both zero-shot and few-shot scenarios, employing metrics such as Accuracy, F1-score, BLEU, and ROUGE for performance evaluation. The results reveal that Gemma 2 consistently outperforms other models across nearly all tasks in both learning paradigms, with particularly strong performance in complex reasoning tasks. However, most models struggle with token-level understanding tasks like Named Entity Recognition, highlighting specific challenges in Persian language processing. This study contributes to the growing body of research on multilingual LLMs, providing valuable insights into their performance in Persian and offering a benchmark for future model development.

大模型评测波斯语少样本学习开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。