Hakim是首个性能超越现有模型的波斯语文本嵌入模型,提升8.5%。
Hakim: Farsi Text Embedding Model
- 基于BERT和RetroMAE架构,结合新数据集训练波斯语嵌入
- 在FaMTEB基准上比旧模型高8.5%准确率,检索任务表现突出
- 适合聊天机器人与带历史记忆的检索增强生成系统
近年来,文本嵌入技术显著提升了多语言自然语言理解能力,但波斯语在大规模嵌入研究中仍严重缺失。本文提出Hakim,一种全新的波斯语文本嵌入模型,在FaMTEB基准上相比现有方法性能提升8.5%,超越所有先前的波斯语模型。为此,我们构建了三个新数据集:Corpesia、Pairsia-sup和Pairsia-unsup,以支持有监督和无监督训练。Hakim适用于聊天机器人与检索增强生成(RAG)系统,尤其在需融合消息历史的检索任务中表现优异。我们还提出一个基于BERT的基线模型,其在各类波斯语NLP任务中均保持更高准确率;而基于RetroMAE的模型在文本信息检索中尤为有效。这些贡献为波斯语理解研究建立了新基础。
原文摘要 · Abstract (English)
Recent advancements in text embedding have significantly improved natural language understanding across many languages, yet Persian remains notably underrepresented in large-scale embedding research. In this paper, we present Hakim, a novel state-of-the-art Persian text embedding model that achieves a 8.5% performance improvement over existing approaches on the FaMTEB benchmark, outperforming all previously developed Persian language models. As part of this work, we introduce three new datasets - Corpesia, Pairsia-sup, and Pairsia-unsup - to support supervised and unsupervised training scenarios. Additionally, Hakim is designed for applications in chatbots and retrieval-augmented generation (RAG) systems, particularly addressing retrieval tasks that require incorporating message history within these systems. We also propose a new baseline model built on the BERT architecture. Our language model consistently achieves higher accuracy across various Persian NLP tasks, while the RetroMAE-based model proves particularly effective for textual information retrieval applications. Together, these contributions establish a new foundation for advancing Persian language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。