IBM推出高效嵌入模型,支持多语言检索,性能超越同类开源模型。
Granite Embedding Models
- 基于编码器架构,融合检索预训练与对比微调技术提升效果。
- 12层主模型在企业检索任务中显著优于同规模公开模型。
- 提供6层轻量版,支持商业与科研用途,已开源至Hugging Face。
我们提出Granite嵌入模型系列,是一组面向检索任务的编码器型嵌入模型,涵盖密集检索与稀疏检索架构,并支持英文及多语言。本报告详述了这些高效12层嵌入模型的技术细节,以及其高效的6层蒸馏版本。通过检索导向预训练、对比微调、知识蒸馏和模型融合等技术,模型在内部IBM检索与搜索任务中显著优于同等规模的公开模型,在广泛使用的信息检索基准上表现相当,且训练数据质量适合企业级应用。所有Granite嵌入模型均以Apache 2.0许可证公开发布,可自由用于研究与商业用途,详情见https://huggingface.co/collections/ibm-granite。
原文摘要 · Abstract (English)
We introduce the Granite Embedding models, a family of encoder-based embedding models designed for retrieval tasks, spanning dense-retrieval and sparse retrieval architectures, with both English and Multilingual capabilities. This report provides the technical details of training these highly effective 12 layer embedding models, along with their efficient 6 layer distilled counterparts. Extensive evaluations show that the models, developed with techniques like retrieval oriented pretraining, contrastive finetuning, knowledge distillation, and model merging significantly outperform publicly available models of similar sizes on both internal IBM retrieval and search tasks, and have equivalent performance on widely used information retrieval benchmarks, while being trained on high-quality data suitable for enterprise use. We publicly release all our Granite Embedding models under the Apache 2.0 license, allowing both research and commercial use at https://huggingface.co/collections/ibm-granite.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。