arXiv:2506.05176cs.CL2025-06被引 1.4k

Qwen3 Embedding提升文本嵌入与重排序,支持多语言和多种部署场景。

Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

  • 基于Qwen3大模型,采用多阶段训练合成高质量多语言数据。
  • 在MTEB等基准上达当前最佳,覆盖代码、跨语言等检索任务。
  • 提供0.6B/4B/8B三尺寸模型,适合不同效率与效果需求。

本文介绍Qwen3 Embedding系列,相较前代GTE-Qwen在文本嵌入与重排序能力上实现显著提升,基于Qwen3基础模型构建。利用Qwen3在多语言理解与生成方面的强大能力,创新设计多阶段训练流程,结合大规模无监督预训练与高质量数据集上的有监督微调。有效的模型融合策略确保了模型的鲁棒性与适应性。训练过程中,Qwen3大模型不仅作为骨干网络,还用于跨领域、多语言合成高质量、丰富且多样化的训练数据,进一步优化训练流程。Qwen3 Embedding系列提供0.6B、4B、8B三种模型规模,适用于嵌入与重排序任务,满足不同部署场景下对效率与效果的权衡。实证评估显示,该系列在多个基准测试中表现优异,尤其在多语言评估基准MTEB上领先,同时在代码检索、跨语言检索及多语言检索任务中均表现出色。为促进可复现性与社区协作,Qwen3 Embedding模型已公开发布,采用Apache 2.0许可证。

原文摘要 · Abstract (English)

In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the Qwen3 LLMs serve not only as backbone models but also play a crucial role in synthesizing high-quality, rich, and diverse training data across multiple domains and languages, thus enhancing the training pipeline. The Qwen3 Embedding series offers a spectrum of model sizes (0.6B, 4B, 8B) for both embedding and reranking tasks, addressing diverse deployment scenarios where users can optimize for either efficiency or effectiveness. Empirical evaluations demonstrate that the Qwen3 Embedding series achieves state-of-the-art results across diverse benchmarks. Notably, it excels on the multilingual evaluation benchmark MTEB for text embedding, as well as in various retrieval tasks, including code retrieval, cross-lingual retrieval and multilingual retrieval. To facilitate reproducibility and promote community-driven research and development, the Qwen3 Embedding models are publicly available under the Apache 2.0 license.

文本嵌入多语言检索大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。