arXiv:2507.14688cs.CLcs.AI2025-07综述

梳理阿拉伯语后训练数据集短板,推动中文读者理解其发展瓶颈

Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations

  • 按能力、可控性、对齐性、鲁棒性四维度评估阿拉伯语数据集
  • 发现任务类型少、文档缺失、社区采用率低等关键缺陷
  • 适合关注中东语言模型发展的研究者与开发者参考

后训练已成为使预训练大语言模型(LLMs)与人类指令对齐的关键技术,显著提升其在各类任务中的表现。该过程的成功依赖于后训练数据集的质量与多样性。本文系统回顾了 Hugging Face Hub 上公开的阿拉伯语后训练数据集,从四大维度展开:(1) 模型能力(如问答、翻译、推理、摘要、对话、代码生成、函数调用);(2) 可控性(如人物设定与系统提示);(3) 对齐性(如文化、安全、伦理与公平);(4) 鲁棒性。每项数据集均依据流行度、实际采纳率、更新频率与维护状态、文档与标注质量、许可证透明度及科学贡献进行严格评估。结果显示,阿拉伯语后训练数据集存在任务多样性不足、文档或标注不一致甚至缺失、社区采纳率低等严重缺口。论文最后讨论这些缺陷对阿拉伯语专用大模型及应用发展的制约,并提出具体改进建议。

原文摘要 · Abstract (English)

Post-training has emerged as a crucial technique for aligning pre-trained Large Language Models (LLMs) with human instructions, significantly enhancing their performance across a wide range of tasks. Central to this process is the quality and diversity of post-training datasets. This paper presents a review of publicly available Arabic post-training datasets on the Hugging Face Hub, organized along four key dimensions: (1) LLM Capabilities (e.g., Question Answering, Translation, Reasoning, Summarization, Dialogue, Code Generation, and Function Calling); (2) Steerability (e.g., Persona and System Prompts); (3) Alignment (e.g., Cultural, Safety, Ethics, and Fairness); and (4) Robustness. Each dataset is rigorously evaluated based on popularity, practical adoption, recency and maintenance, documentation and annotation quality, licensing transparency, and scientific contribution. Our review revealed critical gaps in the development of Arabic post-training datasets, including limited task diversity, inconsistent or missing documentation and annotation, and low adoption across the community. Finally, the paper discusses the implications of these gaps on the progress of Arabic-centric LLMs and applications while providing concrete recommendations for future efforts in Arabic post-training dataset development.

大模型阿拉伯语数据集后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。