arXiv:2511.08376cs.CLcs.IR2025-11被引 3

TurkEmbed提升土耳其语语义理解,专攻推理与相似度任务

TurkEmbed: Turkish Embedding Model on NLI & STS Tasks

  • 融合多源数据与马特罗什卡学习,增强语义表征能力
  • 在土耳其语相似度任务上相关系数提升1-4%,超越现有最佳模型
  • 适合需要高效精准土耳其语理解的NLP应用开发者

本文提出TurkEmbed,一种新型土耳其语嵌入模型,旨在超越现有模型在自然语言推理(NLI)和语义文本相似性(STS)任务中的表现。当前土耳其语嵌入模型多依赖机器翻译数据集,可能影响其准确性和语义理解能力。TurkEmbed结合多样数据集与先进训练技术,包括马特罗什卡表示学习,实现更鲁棒、更精确的嵌入表征。该方法使模型能在资源受限环境下快速编码。在土耳其语STS-b-TR数据集上的评估显示,使用皮尔逊与斯皮尔曼相关系数指标,语义相似性任务性能显著提升。此外,TurkEmbed在All-NLI-TR与STS-b-TR基准测试中均超越当前最优模型Emrecan,提升幅度达1-4%。TurkEmbed有望推动土耳其语NLP生态发展,深化语言理解,助力下游应用进步。

原文摘要 · Abstract (English)

This paper introduces TurkEmbed, a novel Turkish language embedding model designed to outperform existing models, particularly in Natural Language Inference (NLI) and Semantic Textual Similarity (STS) tasks. Current Turkish embedding models often rely on machine-translated datasets, potentially limiting their accuracy and semantic understanding. TurkEmbed utilizes a combination of diverse datasets and advanced training techniques, including matryoshka representation learning, to achieve more robust and accurate embeddings. This approach enables the model to adapt to various resource-constrained environments, offering faster encoding capabilities. Our evaluation on the Turkish STS-b-TR dataset, using Pearson and Spearman correlation metrics, demonstrates significant improvements in semantic similarity tasks. Furthermore, TurkEmbed surpasses the current state-of-the-art model, Emrecan, on All-NLI-TR and STS-b-TR benchmarks, achieving a 1-4\% improvement. TurkEmbed promises to enhance the Turkish NLP ecosystem by providing a more nuanced understanding of language and facilitating advancements in downstream applications.

土耳其语语义嵌入自然语言推理相似度计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。