构建首个面向阿拉伯语大模型的综合性评测平台,解决数据匮乏与评估不透明问题。
BALSAM: A Platform for Benchmarking Arabic Large Language Models
- 设计覆盖14类任务的78项NLP评测,含52000个样本
- 提供37000个测试集样本与15000个开发集样本,支持盲测
- 社区驱动平台,推动阿拉伯语大模型协作研究
英语大语言模型的迅猛进展并未在所有语言中同步。阿拉伯语大模型性能滞后,主要因数据稀缺、语言多样性、方言差异及形态复杂性等挑战。现有评测质量不足,普遍依赖静态公开数据,任务覆盖不全,且缺乏独立测试集和统一平台,难以衡量真实进步并防止数据泄露。为此,我们提出BALSAM——一个面向阿拉伯语大模型的综合性、社区共建评测平台。该平台涵盖14个类别中的78项自然语言处理任务,共52,000个样本(37,000个测试集,15,000个开发集),并提供集中化、透明化的盲测环境。BALSAM旨在成为统一标准,促进协作研究,推动阿拉伯语大模型能力提升。
原文摘要 · Abstract (English)
The impressive advancement of Large Language Models (LLMs) in English has not been matched across all languages. In particular, LLM performance in Arabic lags behind, due to data scarcity, linguistic diversity of Arabic and its dialects, morphological complexity, etc. Progress is further hindered by the quality of Arabic benchmarks, which typically rely on static, publicly available data, lack comprehensive task coverage, or do not provide dedicated platforms with blind test sets. This makes it challenging to measure actual progress and to mitigate data contamination. Here, we aim to bridge these gaps. In particular, we introduce BALSAM, a comprehensive, community-driven benchmark aimed at advancing Arabic LLM development and evaluation. It includes 78 NLP tasks from 14 broad categories, with 52K examples divided into 37K test and 15K development, and a centralized, transparent platform for blind evaluation. We envision BALSAM as a unifying platform that sets standards and promotes collaborative research to advance Arabic LLM capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。