arXiv:2411.01192cs.CL2024-11NAACL被引 16

推出专精阿拉伯语的嵌入模型与评测基准,支持多方言多文化场景。

Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks

  • 基于ARBERTv2和ArMistral构建小型与大型阿拉伯语嵌入模型。
  • 在94个数据集上覆盖8项任务,大模型性能超越Multilingual-E5-large。
  • 兼顾方言与文化敏感性,适合阿拉伯语NLP研究与实际应用者使用。

我们提出{f Swan}系列嵌入模型,聚焦阿拉伯语,覆盖小规模与大规模应用场景。其中,Swan-Small基于ARBERTv2,Swan-Large则基于预训练阿拉伯语大模型ArMistral。为评估模型性能,我们构建了ArabicMTEB基准测试套件,涵盖跨语言、多方言、多领域及跨文化任务,共包含8类任务和94个数据集。Swan-Large在多数阿拉伯语任务中表现优于Multilingual-E5-large,而Swan-Small始终超越Multilingual-E5-base。大量实验表明,Swan模型具备方言与文化感知能力,在多种阿拉伯语场景下表现优异,并具有显著的成本效益。本工作极大推进了阿拉伯语建模研究,为未来阿拉伯语自然语言处理提供重要资源。模型与基准代码已开源:https://github.com/UBC-NLP/swan

原文摘要 · Abstract (English)

We introduce {\bf Swan}, a family of embedding models centred around the Arabic language, addressing both small-scale and large-scale use cases. Swan includes two variants: Swan-Small, based on ARBERTv2, and Swan-Large, built on ArMistral, a pretrained Arabic large language model. To evaluate these models, we propose ArabicMTEB, a comprehensive benchmark suite that assesses cross-lingual, multi-dialectal, multi-domain, and multi-cultural Arabic text embedding performance, covering eight diverse tasks and spanning 94 datasets. Swan-Large achieves state-of-the-art results, outperforming Multilingual-E5-large in most Arabic tasks, while the Swan-Small consistently surpasses Multilingual-E5-base. Our extensive evaluations demonstrate that Swan models are both dialectally and culturally aware, excelling across various Arabic domains while offering significant monetary efficiency. This work significantly advances the field of Arabic language modelling and provides valuable resources for future research and applications in Arabic natural language processing. Our models and benchmark are available at our GitHub page: \href{https://github.com/UBC-NLP/swan}{https://github.com/UBC-NLP/swan}

阿拉伯语嵌入模型多方言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。