专为土耳其语检索优化的嵌入模型,性能超越现有方法。
TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- 基于土耳其语语义理解模型,通过多重负样本损失微调
- 在Scifact TR数据集上比colBERT高出19.26%的检索准确率
- 适合需要高精度土耳其语信息检索的研究与应用
本文提出TurkEmbed4Retrieval,是原用于自然语言推理和语义文本相似度任务的TurkEmbed模型的检索专用版本。通过在MS MARCO TR数据集上使用马特里什卡表示学习和定制化的多负样本排序损失进行微调,该模型在土耳其语检索任务中达到最先进水平。大量实验表明,在Scifact TR数据集的关键检索指标上,其性能相比土耳其语colBERT提升19.26%,确立了新的基准。
原文摘要 · Abstract (English)
In this work, we introduce TurkEmbed4Retrieval, a retrieval specialized variant of the TurkEmbed model originally designed for Natural Language Inference (NLI) and Semantic Textual Similarity (STS) tasks. By fine-tuning the base model on the MS MARCO TR dataset using advanced training techniques, including Matryoshka representation learning and a tailored multiple negatives ranking loss, we achieve SOTA performance for Turkish retrieval tasks. Extensive experiments demonstrate that our model outperforms Turkish colBERT by 19,26% on key retrieval metrics for the Scifact TR dataset, thereby establishing a new benchmark for Turkish information retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。