arXiv:2608.24477cs.CL2026-08

揭示斯拉夫语嵌入模型评估受限于数据稀缺,提出新框架提升评估可靠性。

Dataset Scarcity Limits Robust Evaluation of Multilingual Embedding Models: A Case Study of Slavic Languages

论文配图:Dataset Scarcity Limits Robust Evaluation of Multilingual Embedding Models: A Case Study of Slavic Languages
图 1 · 摘自论文原文
  • 构建二维评估框架,区分任务内与跨任务分析
  • 发现多数斯拉夫语任务仅依赖单一数据集,结论不可靠
  • 识别出llama-embed-nemotron-8b等高泛化能力模型

多语言文本嵌入模型支持多种自然语言处理任务的跨语言知识迁移,但其评估在高、中、低资源语言间极不均衡。本文针对多语言嵌入基准在数据稀缺下的问题,提出一个二维分析框架,并应用于MTEB基准中的斯拉夫语子集。该框架区分任务特定与跨任务评估,联合分析三个维度:(1)排名鲁棒性,(2)模型一致性,(3)证据强度。任务特定层面评估模型排名在排序方法和基准数据集组成变化下的稳定性;跨任务层面评估模型在单语种内多种任务间的泛化能力。为量化结论可靠性,引入证据强度评分,综合考虑数据可用性、多样性及可鲁棒评估性。分析显示存在严重基准稀疏,许多斯拉夫语言-任务对仅依赖单一数据集或高度相关的基准集合,限制了稳健结论的得出。跨任务分析发现少数模型具有强泛化能力,尤其包括llama-embed-nemotron-8b、multilingual-e5-large-instruct和Qwen3-Embedding变体,在多种斯拉夫语言和任务上表现稳定。整体表明,基准排名与鲁棒性结论必须结合证据强度理解,数据稀缺是可信多语言评估的主要障碍。

原文摘要 · Abstract (English)

Multilingual text embedding models enable cross-lingual transfer of knowledge across a wide range of NLP tasks, but their evaluation remains highly uneven across high-, mid- and low-resource languages. In this paper, we propose a two-dimensional framework, specifically tailored for analyzing multilingual embedding benchmarks under dataset scarcity, and apply it on the Slavic-language subset of the MTEB benchmark. The framework distinguishes between task-specific and cross-task evaluation, while jointly analyzing three complementary aspects: (1) ranking robustness, (2) model consistency, and (3) evidence strength. At the task-specific level, we evaluate the stability of model rankings under changes in ranking methodology and benchmark dataset composition. At the cross-task level, we assess the ability of models to generalize across diverse tasks within a language. To quantify the reliability of benchmark conclusions, we introduce an Evidence Strength Score that accounts for dataset availability, diversity, and robustness assessability. Our analysis reveals severe benchmark sparsity, with many Slavic language-task pairs relying on a single dataset or highly correlated benchmark collections, limiting the ability to draw robust conclusions. The cross-task analysis reveals a small group of highly transferable models, most notably llama-embed-nemotron-8b, multilingual-e5-large-instruct, and Qwen3-Embedding variants, that consistently perform well across Slavic languages and tasks. Overall, the results demonstrate that benchmark rankings and robustness conclusions must be interpreted jointly with certain notation of their evidence strength and highlight benchmark scarcity as a major obstacle to trustworthy multilingual evaluation.

多语言嵌入评估框架数据稀缺斯拉夫语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。