arXiv:2501.16516cs.CLcs.AI2025-01被引 14

评测大模型对阿拉伯语作文的评分能力,发现提示工程显著影响效果。

How well can LLMs Grade Essays in Arabic?

  • 测试多种大模型在阿拉伯语作文评分中的表现,对比零样本、少样本与微调方法。
  • ACEGPT在数据集上达QWK 0.67,但小于小型BERT模型的0.88。
  • 混合语言提示提升理解,适合教育评估领域研究者参考。

本研究评估了当前最先进的大语言模型(包括ChatGPT、Llama、Aya、Jais和ACEGPT)在阿拉伯语自动作文评分(AR-AES)任务中的表现,使用真实学生数据的AR-AES数据集。研究探索了零样本、少样本上下文学习及微调等多种评估方法,并通过在提示中加入评分标准来考察指令遵循能力的影响。采用中英混合提示策略以增强模型对阿拉伯语文本的理解与表现。测试结果显示,ACEGPT在整体表现上最优,达到二次加权卡帕系数(QWK)0.67,但仍低于一个小型BERT模型的0.88。研究揭示了大模型处理阿拉伯语时面临分词复杂性和更高计算开销等挑战。不同课程间性能差异表明,需开发能适应多样化评估格式的自适应模型。结果表明有效提示工程可显著提升大模型输出质量。据我们所知,这是首个基于真实学生数据,系统评估多个生成式大模型在阿拉伯语作文评分任务上的实证研究。

原文摘要 · Abstract (English)

This research assesses the effectiveness of state-of-the-art large language models (LLMs), including ChatGPT, Llama, Aya, Jais, and ACEGPT, in the task of Arabic automated essay scoring (AES) using the AR-AES dataset. It explores various evaluation methodologies, including zero-shot, few-shot in-context learning, and fine-tuning, and examines the influence of instruction-following capabilities through the inclusion of marking guidelines within the prompts. A mixed-language prompting strategy, integrating English prompts with Arabic content, was implemented to improve model comprehension and performance. Among the models tested, ACEGPT demonstrated the strongest performance across the dataset, achieving a Quadratic Weighted Kappa (QWK) of 0.67, but was outperformed by a smaller BERT-based model with a QWK of 0.88. The study identifies challenges faced by LLMs in processing Arabic, including tokenization complexities and higher computational demands. Performance variation across different courses underscores the need for adaptive models capable of handling diverse assessment formats and highlights the positive impact of effective prompt engineering on improving LLM outputs. To the best of our knowledge, this study is the first to empirically evaluate the performance of multiple generative Large Language Models (LLMs) on Arabic essays using authentic student data.

阿拉伯语作文评分大模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。