IBM推出高性能嵌入模型,支持超长文本检索与企业级部署。
Granite Embedding R2 Models
- 基于双编码器与交叉编码器架构,上下文长度达8192词元。
- 在文本、代码、长文档等多领域表现领先,速度比竞品快19%-44%。
- 开源免费,适合企业级应用与研究,数据来源透明可追溯。
我们推出Granite Embedding R2系列模型,一套专为企业级密集检索场景设计的高性能英文编码器嵌入模型。基于首代版本,该系列模型实现显著提升:上下文长度扩展至8,192词元(提升16倍),在文本、代码、长文档搜索、多轮对话及表格数据等多种检索任务中达到业界领先水平,并在保持高精度的同时,相较主流竞品提速19%-44%。模型包含22层高效检索器及其12层轻量版,以及高质量重排序模型,均基于企业适用数据并经过全面治理训练。在标准基准、IBM自研评估套件及真实企业场景中表现优异,确立了开源嵌入模型的新标准。所有模型均以Apache 2.0许可证公开,可在Hugging Face平台自由使用,支持科研与商业应用。
原文摘要 · Abstract (English)
We introduce the Granite Embedding R2 models, a comprehensive family of high-performance English encoder-based embedding models engineered for enterprise-scale dense retrieval applications. Building upon our first-generation release, these models deliver substantial improvements, including 16x expanded context length (8,192 tokens), state-of-the-art performance across diverse retrieval domains - text, code, long-document search, multi-turn conversational, and tabular data - and measurable speed advantages of 19-44\% over leading competitors while maintaining superior accuracy. Our release encompasses both bi-encoder and cross-encoder architectures, featuring a highly effective 22-layer retriever model and its efficient 12-layer counterpart, alongside a high-quality reranker model, all trained exclusively on enterprise-appropriate data with comprehensive governance oversight. The models demonstrate exceptional versatility across standard benchmarks, IBM-developed evaluation suites, and real-world enterprise use cases, establishing new performance standards for open-source embedding models. In an era where retrieval speed and accuracy are paramount for competitive advantage, the Granite R2 models deliver a compelling combination of cutting-edge performance, enterprise-ready licensing, and transparent data provenance that organizations require for mission-critical deployments. All models are publicly available under the Apache 2.0 license at https://huggingface.co/collections/ibm-granite, enabling unrestricted research and commercial use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。