arXiv:2606.21460cs.CLcs.AI2026-06

评测12个小型模型在阿拉伯语任务上的表现,发现语言对齐比模型大小更重要。

Evaluation of Small Language Models for Arabic Language Processing

论文配图:Evaluation of Small Language Models for Arabic Language Processing
图 1 · 摘自论文原文
  • 构建240项跨8域10技能的阿拉伯语评测集,零样本统一提示
  • Gemma 3 (12B)得分最高(4.548/5),性能与语言对齐相关
  • 揭示幻觉、提示泄露等共性缺陷,适合做阿拉伯语AI研究者参考

本文评估了12个小型语言模型(SLMs)在阿拉伯语自然语言处理任务中的表现。研究构建了一个包含240个测试项的基准,覆盖8个领域和10种语言技能,涵盖理解型与生成型任务。所有模型在标准化阿拉伯语提示模板下进行零样本评估。模型输出由GPT-4.1 Mini、Claude Haiku 4.5和DeepSeek-Chat组成的多模型评分框架打分,结果按任务、技能和模型家族分析。结果显示,Gemma 3 (12B)总体得分最高(4.548/5),其次为Aya和C4AI Command Arabic。研究发现模型规模并非决定性能的关键,具备更强阿拉伯语对齐和更可靠指令遵循能力的模型表现更优。低分模型普遍存在提示泄漏、幻觉、语言漂移、生成不完整及任务偏离等问题。该基准为紧凑型阿拉伯语模型评估提供结构化参考,助力高效、可靠且文化适配的阿拉伯语AI系统发展。

原文摘要 · Abstract (English)

This paper evaluates the performance of twelve Small Language Models (SLMs) on Arabic natural language processing tasks. The study introduces a benchmark of 240 Arabic test items distributed across eight domains and ten language skills, covering both comprehension-oriented and generation-oriented tasks. All models were evaluated under a controlled zero-shot setting using a standardized Arabic-only prompt template. Model responses were assessed through a multi-model LLM-as-a-judge framework involving GPT-4.1 Mini, Claude Haiku 4.5, and DeepSeek-Chat, with scores aggregated across judges and analyzed by task, skill, and model family. The results show that Gemma 3 (12B) achieved the highest overall score (4.548/5), followed by Aya and C4AI Command Arabic. The observed results suggest that model size alone does not explain Arabic SLM performance. Models with stronger Arabic alignment and more reliable instruction-following behavior tended to perform better across tasks. Common failure patterns among lower-performing models include prompt leakage, hallucination, language drift, incomplete generation, and weak task adherence. Overall, the benchmark provides a structured reference for evaluating compact Arabic language models and supports future work on efficient, reliable, and culturally appropriate Arabic AI systems.

小模型阿拉伯语评测基准语言对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。