用结构化提示提升阿拉伯语作文评分,精准评估组织、词汇等语言能力。
Structured Prompting for Arabic Essay Proficiency: A Trait-Centric Evaluation Approach
- 设计三层次提示策略,模拟专家评分员分项评估作文
- Fanar-1-9B-Instruct在零样本和少样本下表现最佳,最高QWK达0.28
- 基于评分标准的提示显著提升论述与风格等高阶能力评分
本文提出一种面向阿拉伯语作文能力的新型提示工程框架,利用大语言模型在零样本和少样本条件下实现特质化自动作文评分。针对阿拉伯语缺乏可扩展且语言学导向的自动评分工具的问题,引入三层提示策略(标准、混合、基于评分量表),引导模型评估组织、词汇、展开和风格等语言能力特质。混合方法模拟多代理评分机制,基于评分量表的方法则通过已评分样例增强模型对齐。在首个公开的带特质标注的阿拉伯语作文数据集QAES上,评估了八种大语言模型。结果表明,在零样本和少样本设置下,Fanar-1-9B-Instruct在所有特质上达到最高一致性(QWK=0.28,置信区间=0.41),基于评分量表的提示在所有模型和特质中均带来稳定提升,尤其在论述与风格等话语层面特质改善显著。研究证实,结构化提示比模型规模更能有效实现阿拉伯语自动作文评分。本工作首次构建了以能力为导向的阿拉伯语自动评分综合框架,为低资源教育场景下的可扩展评估奠定基础。
原文摘要 · Abstract (English)
This paper presents a novel prompt engineering framework for trait specific Automatic Essay Scoring (AES) in Arabic, leveraging large language models (LLMs) under zero-shot and few-shot configurations. Addressing the scarcity of scalable, linguistically informed AES tools for Arabic, we introduce a three-tier prompting strategy (standard, hybrid, and rubric-guided) that guides LLMs in evaluating distinct language proficiency traits such as organization, vocabulary, development, and style. The hybrid approach simulates multi-agent evaluation with trait specialist raters, while the rubric-guided method incorporates scored exemplars to enhance model alignment. In zero and few-shot settings, we evaluate eight LLMs on the QAES dataset, the first publicly available Arabic AES resource with trait level annotations. Experimental results using Quadratic Weighted Kappa (QWK) and Confidence Intervals show that Fanar-1-9B-Instruct achieves the highest trait level agreement in both zero and few-shot prompting (QWK = 0.28 and CI = 0.41), with rubric-guided prompting yielding consistent gains across all traits and models. Discourse-level traits such as Development and Style showed the greatest improvements. These findings confirm that structured prompting, not model scale alone, enables effective AES in Arabic. Our study presents the first comprehensive framework for proficiency oriented Arabic AES and sets the foundation for scalable assessment in low resource educational contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。