首个面向波斯语的文本嵌入综合评测基准,覆盖63个数据集。
FaMTEB: Massive Text Embedding Benchmark in Persian Language
- 构建波斯语文本嵌入评测基准,整合现有、翻译与新生成数据。
- 涵盖7类任务,包含首次加入的聊天机器人与摘要检索任务。
- 开源数据集、代码与排行榜,助力波斯语模型评估与训练。
本文提出一个面向波斯语(波斯文)的全面文本嵌入评测基准,基于大规模文本嵌入基准(MTEB)。该基准包含63个数据集,覆盖分类、聚类、成对分类、重排序、检索、摘要检索和语义文本相似性共七类任务。数据来源包括已有数据、翻译数据及新生成数据,形成多样化的波斯语模型评估框架。随着文本嵌入模型在聊天机器人和检索增强生成系统中的广泛应用,评测数据集已成为关键组成部分。本工作首次将聊天机器人评测数据集纳入MTEB,并引入标准MTEB中未包含的摘要检索新任务。此外,本文还发布了大量适用于波斯语训练与评估的新数据集,其中部分为波斯语领域首次出现。我们评估了多个波斯语及多语言嵌入模型在各类任务上的表现。本研究提供开源基准,包含数据集、代码与公开排行榜。
原文摘要 · Abstract (English)
In this paper, we introduce a comprehensive benchmark for Persian (Farsi) text embeddings, built upon the Massive Text Embedding Benchmark (MTEB). Our benchmark includes 63 datasets spanning seven different tasks: classification, clustering, pair classification, reranking, retrieval, summary retrieval, and semantic textual similarity. The datasets are formed as a combination of existing, translated, and newly generated data, offering a diverse evaluation framework for Persian language models. Given the increasing use of text embedding models in chatbots, evaluation datasets are becoming inseparable ingredients in chatbot challenges and Retrieval-Augmented Generation systems. As a contribution, we include chatbot evaluation datasets in the MTEB benchmark for the first time. In addition, in this paper, we introduce the new task of summary retrieval which is not part of the tasks included in standard MTEB. Another contribution of this paper is the introduction of a substantial number of new Persian language NLP datasets suitable for training and evaluation, some of which have no previous counterparts in Persian. We evaluate the performance of several Persian and multilingual embedding models in a range of tasks. This work introduces an open-source benchmark with datasets, code and a public leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。