填补荷兰语嵌入模型空白,推出首个专为荷兰语设计的评测基准与高效模型。
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
- 构建覆盖多任务的荷兰语嵌入评测基准MTEB-NL,整合现有与新创建数据集。
- 基于真实与大模型生成的合成数据训练出高性能的E5-NL系列模型。
- 资源开源,适合从事荷兰语NLP研究或本地化应用的开发者使用。
近年来,多种语言的嵌入资源(包括模型、基准和数据集)被广泛发布,以支持多语言应用。然而,荷兰语仍严重缺乏代表性,通常仅占多语言资源的一小部分。为弥补这一差距并推动荷兰语嵌入技术的发展,我们推出了新的评估与生成资源。首先,提出面向荷兰语的大型文本嵌入基准MTEB-NL,涵盖现有与新创建的数据集,覆盖广泛任务。其次,构建一个由可用荷兰语检索数据集组成的训练数据集,并结合大语言模型生成的合成数据,扩展任务范围至检索之外。最后,发布一系列紧凑而高效的E5-NL嵌入模型,在多个任务上表现优异。所有资源均已通过Hugging Face Hub与MTEB工具包公开。
原文摘要 · Abstract (English)
Recently, embedding resources, including models, benchmarks, and datasets, have been widely released to support a variety of languages. However, the Dutch language remains underrepresented, typically comprising only a small fraction of the published multilingual resources. To address this gap and encourage the further development of Dutch embeddings, we introduce new resources for their evaluation and generation. First, we introduce the Massive Text Embedding Benchmark for Dutch (MTEB-NL), which includes both existing Dutch datasets and newly created ones, covering a wide range of tasks. Second, we provide a training dataset compiled from available Dutch retrieval datasets, complemented with synthetic data generated by large language models to expand task coverage beyond retrieval. Finally, we release a series of E5-NL models compact yet efficient embedding models that demonstrate strong performance across multiple tasks. We make our resources publicly available through the Hugging Face Hub and the MTEB package.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。