arXiv:2510.23896cs.CL2025-10Conference of the …被引 4

为非洲语言构建首个大规模文本嵌入基准与适配模型

AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages

  • 构建覆盖59种非洲语言的AfriMTEB基准,含14项任务和38个数据集
  • 新任务如仇恨言论检测等首次覆盖,部分数据集涵盖56种语言
  • 基于mE5模型通过对比蒸馏适配非洲语言,性能超越Gemini等基线

文本嵌入是检索增强生成等NLP任务的关键组件,对防止大模型幻觉至关重要。尽管已有大规模多语言MTEB(MMTEB)发布,非洲语言仍严重不足,现有任务多复用翻译基准(如FLORES聚类或SIB-200)。本文提出AfriMTEB——MMTEB的区域扩展,涵盖59种语言、14项任务和38个数据集,新增6个数据集。不同于多数仅含少于5种语言的MMTEB数据集,新数据集覆盖14至56种非洲语言,并引入仇恨言论检测、意图识别和情感分类等全新任务。同时,提出AfriE5,通过跨语言对比蒸馏将指令微调的mE5模型适配非洲语言。评估显示,AfriE5性能达当前最优,优于Gemini-Embeddings和mE5等强基线。

原文摘要 · Abstract (English)

Text embeddings are an essential building component of several NLP tasks such as retrieval-augmented generation which is crucial for preventing hallucinations in LLMs. Despite the recent release of massively multilingual MTEB (MMTEB), African languages remain underrepresented, with existing tasks often repurposed from translation benchmarks such as FLORES clustering or SIB-200. In this paper, we introduce AfriMTEB -- a regional expansion of MMTEB covering 59 languages, 14 tasks, and 38 datasets, including six newly added datasets. Unlike many MMTEB datasets that include fewer than five languages, the new additions span 14 to 56 African languages and introduce entirely new tasks, such as hate speech detection, intent detection, and emotion classification, which were not previously covered. Complementing this, we present AfriE5, an adaptation of the instruction-tuned mE5 model to African languages through cross-lingual contrastive distillation. Our evaluation shows that AfriE5 achieves state-of-the-art performance, outperforming strong baselines such as Gemini-Embeddings and mE5.

文本嵌入非洲语言基准测试模型适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。