arXiv:2603.19223cs.CLcs.AI2026-03被引 9

F2LLM-v2打造多语言高效嵌入模型,覆盖200+语言且小模型表现领先。

F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World

  • 两阶段大模型训练结合分层学习与剪枝,提升效率
  • 14B模型在11项MTEB基准上排名第一,小模型也创资源受限新纪录
  • 开源全部模型数据代码,助力多语言研究

我们提出F2LLM-v2,一个涵盖8种规模(80M至14B)的通用多语言嵌入模型家族。基于新整理的6000万条高质量公开数据样本训练,支持超过200种语言,尤其关注中低资源语言。通过两阶段LLM嵌入训练流程,结合马特罗什卡学习、模型剪枝和知识蒸馏技术,实现比以往基于大模型的嵌入模型更高效的性能。大量评估显示,F2LLM-v2-14B在11个MTEB基准中排名第一,家族中的小型模型也在资源受限场景中创下新SOTA。为推动开源嵌入模型研究,我们发布了所有模型、数据、代码及中间检查点。

原文摘要 · Abstract (English)

We present F2LLM-v2, a new family of general-purpose, multilingual embedding models in 8 distinct sizes ranging from 80M to 14B. Trained on a newly curated composite of 60 million publicly available high-quality data samples, F2LLM-v2 supports more than 200 languages, with a particular emphasis on previously underserved mid- and low-resource languages. By integrating a two-stage LLM-based embedding training pipeline with matryoshka learning, model pruning, and knowledge distillation techniques, we present models that are far more efficient than previous LLM-based embedding models while retaining competitive performances. Extensive evaluations confirm that F2LLM-v2-14B ranks first on 11 MTEB benchmarks, while the smaller models in the family also set a new state of the art for resource-constrained applications. To facilitate open-source embedding model research, we release all models, data, code, and intermediate checkpoints.

多语言嵌入大模型高效模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。