arXiv:2505.24581cs.CL2025-05被引 12

GATE模型提升阿拉伯语语义相似度,性能超越大模型20%-25%。

GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training

  • 采用马特里什卡表征学习与混合损失训练,增强细粒度语义理解。
  • 在MTEB基准上实现领先性能,比大型模型提升20%-25%。
  • 专为阿拉伯语设计,适合需要精准语义匹配的应用场景。

语义文本相似度(STS)是自然语言处理中的关键任务,广泛应用于信息检索、聚类及文本间语义关系理解。然而,由于高质量数据集和预训练模型的缺乏,阿拉伯语领域的相关研究仍显不足,限制了语义相似度技术的准确评估与进展。本文提出通用阿拉伯语文本嵌入模型GATE,其在MTEB基准上的语义文本相似度任务中达到当前最佳表现。GATE结合马特里什卡表征学习与基于阿拉伯语三元组数据集的自然语言推断任务的混合损失训练方法,显著提升模型在需精细语义理解任务中的性能。实验表明,GATE在多个STS基准上相较更大规模模型(如OpenAI)实现20%-25%的性能提升,有效捕捉阿拉伯语的独特语义特征。

原文摘要 · Abstract (English)

Semantic textual similarity (STS) is a critical task in natural language processing (NLP), enabling applications in retrieval, clustering, and understanding semantic relationships between texts. However, research in this area for the Arabic language remains limited due to the lack of high-quality datasets and pre-trained models. This scarcity of resources has restricted the accurate evaluation and advance of semantic similarity in Arabic text. This paper introduces General Arabic Text Embedding (GATE) models that achieve state-of-the-art performance on the Semantic Textual Similarity task within the MTEB benchmark. GATE leverages Matryoshka Representation Learning and a hybrid loss training approach with Arabic triplet datasets for Natural Language Inference, which are essential for enhancing model performance in tasks that demand fine-grained semantic understanding. GATE outperforms larger models, including OpenAI, with a 20-25% performance improvement on STS benchmarks, effectively capturing the unique semantic nuances of Arabic.

阿拉伯语文本嵌入语义相似度表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。