高效训练俄语文本嵌入模型,性能超越现有方法
GigaEmbeddings: Efficient Russian Language Embedding Model
- 三阶段训练:大规模对比预训练、硬负样本微调、多任务泛化
- 在ruMTEB基准上达69.1分,参数更多仍更高效
- 创新设计提升上下文建模与序列聚合,剪枝25%层不降性能
我们提出GigaEmbeddings,一种通过分层指令微调专为俄语设计的解码器仅用大型语言模型(GigaChat-3B)训练高性能俄语文本嵌入的新框架。其三阶段流程包括:在网页规模语料库中进行大规模对比预训练、使用硬负样本微调,以及在检索、分类和聚类任务间的多任务泛化,通过统一多样化目标并利用合成数据生成,解决现有方法的关键局限。架构创新包括双向注意力用于上下文建模、潜在注意力池化实现鲁棒序列聚合,以及战略性地剪枝25%的Transformer层以提升效率而不牺牲性能。在涵盖23个跨语言任务的ruMTEB基准上评估,GigaEmbeddings取得69.1的平均得分,优于参数更多但表现较差的强基线。
原文摘要 · Abstract (English)
We introduce GigaEmbeddings, a novel framework for training high-performance Russian-focused text embeddings through hierarchical instruction tuning of the decoder-only LLM designed specifically for Russian language (GigaChat-3B). Our three-stage pipeline, comprising large-scale contrastive pre-training in web-scale corpora, fine-tuning with hard negatives, and multitask generalization across retrieval, classification, and clustering tasks, addresses key limitations of existing methods by unifying diverse objectives and leveraging synthetic data generation. Architectural innovations include bidirectional attention for contextual modeling, latent attention pooling for robust sequence aggregation, and strategic pruning of 25% of transformer layers to enhance efficiency without compromising performance. Evaluated on the ruMTEB benchmark spanning 23 multilingual tasks, GigaEmbeddings achieves state-of-the-art results (69.1 avg. score), outperforming strong baselines with a larger number of parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。