arXiv:2501.01028cs.CL2025-01被引 48

用高质量数据训练出更强的多语言嵌入模型

KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model

  • 用大模型生成多样化合成数据,提升训练多样性
  • 在多语言MTEB基准上超越同类小模型,性能新高
  • 适合需要精准语义匹配的跨语言应用开发者

随着检索增强生成在大模型中的普及,嵌入模型的重要性日益凸显。尽管通用嵌入模型数量增加,但以往研究常忽视训练数据质量的关键作用。本文提出KaLM-Embedding,一个基于大量更清洁、更多样、更具领域特性的训练数据的通用多语言嵌入模型。采用三项关键技术:(1) 基于角色的合成数据,从大模型中提炼多样化样本;(2) 排名一致性过滤,剔除信息量低的样本;(3) 半同质任务批次采样,提升训练效率。不同于传统BERT类架构,我们选用Qwen2-0.5B作为预训练基础,实现自回归语言模型向通用嵌入任务的适配。在涵盖多种语言的MTEB基准上进行广泛评估,结果表明该模型在参数量小于1亿的情况下,性能优于同等规模其他模型,为<1B参数的多语言嵌入模型设立了新标准。

原文摘要 · Abstract (English)

As retrieval-augmented generation prevails in large language models, embedding models are becoming increasingly crucial. Despite the growing number of general embedding models, prior work often overlooks the critical role of training data quality. In this work, we introduce KaLM-Embedding, a general multilingual embedding model that leverages a large quantity of cleaner, more diverse, and domain-specific training data. Our model has been trained with key techniques proven to enhance performance: (1) persona-based synthetic data to create diversified examples distilled from LLMs, (2) ranking consistency filtering to remove less informative samples, and (3) semi-homogeneous task batch sampling to improve training efficacy. Departing from traditional BERT-like architectures, we adopt Qwen2-0.5B as the pre-trained model, facilitating the adaptation of auto-regressive language models for general embedding tasks. Extensive evaluations of the MTEB benchmark across multiple languages show that our model outperforms others of comparable size, setting a new standard for multilingual embedding models with <1B parameters.

嵌入模型多语言数据质量大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。