评估多语言文本嵌入模型在不同任务和数据下的稳定性表现
On the Robustness of Multilingual Text Embedding Rankings Across Learning Tasks, Languages, and Benchmark Datasets

- 用多种决策方法分析模型排名变化,检验结果可靠性
- 发现大模型在多数任务中表现稳定,但检索任务例外
- 适合关注跨语言评估公平性的研究人员和开发者
大规模多语言文本嵌入模型在科研与工业中至关重要,但其在特定语言和多任务场景下的表现仍不清晰。尽管MTEB平台覆盖超过250种语言,但模型优劣结论常依赖于隐含的数据集构成与聚合方式选择。为此,我们对MTEB中的多语言模型性能进行元研究,采用多样化的多准则决策排名方法,并引入两个鲁棒性指标:数据集构成鲁棒性(排名对数据集组成变化的敏感度)与排名方案鲁棒性(对聚合方法变化的敏感度),实现对基准测试结论稳定性的系统性敏感性分析。我们在英语、法语、德语、印地语和西班牙语五种语言上,针对九类任务(如分类、聚类、检索)开展深入分析,并公开约230种额外语言的结果。任务特定分析显示,基于大语言模型的模型通常为稳健领先者,但并非完全一致(如在检索任务中);而任务无关结果表明,仅有极少数模型能在各类任务、排名方案和数据子样本中持续保持优异表现。
原文摘要 · Abstract (English)
Large-scale multilingual text embedding models play crucial role in both research and industry, yet their behavior in language-specific, multi-task settings remains insufficiently understood. Although benchmarking platforms such as MTEB report results across more than 250 languages, conclusions about model superiority often depend on implicit choices of dataset compositions and performance aggregation methods. To address this gap, we present a meta-study of multilingual model performance robustness in MTEB, applying a diverse set of multi-criteria decision-making ranking schemes and introducing two robustness indicators: dataset-composition robustness (sensitivity of rankings to changing dataset compositions) and ranking-scheme robustness (sensitivity to aggregation method change). They enable systematic sensitivity analysis of whether benchmarking conclusions remain stable under different evaluation designs. We conduct an in-depth analysis on five languages (English, French, German, Hindi, and Spanish) across nine tasks (e.g., classification, clustering, retrieval) and release results for approximately 230 additional languages. The task-specific analyses show that large-scale LLM-based models are often robust top performers, though not uniformly (e.g., in retrieval task), while task-agnostic results reveal that only a small subset of models remains consistently strong across tasks, ranking schemes, and data subsamples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。