对比7种模型在捷克语检索上的表现,找最优方案
A Comparative Study of Text Retrieval Models on DaReCzech
- 用7个现成模型直接或翻译后检索捷克语数据
- Gemma22表现最佳,Contriever最差,SPLADE和PLAID平衡好
- 适合需要高效精准捷克语检索的研究者
本文对7个现成文档检索模型(Splade、Plaid、Plaid-X、SimCSE、Contriever、OpenAI ADA、Gemma2)在捷克语数据集DaReCzech上的表现进行了全面评估。目标是衡量现代检索方法在捷克语中的性能。实验涵盖检索质量、速度和内存占用,并分析了直接使用捷克语文本与先翻译成英语再检索的优劣。结果表明各模型表现差异显著:Gemma22在精度和召回率上最高,Contriever表现最差;SPLADE和PLAID在效率与性能间取得良好平衡。
原文摘要 · Abstract (English)
This article presents a comprehensive evaluation of 7 off-the-shelf document retrieval models: Splade, Plaid, Plaid-X, SimCSE, Contriever, OpenAI ADA and Gemma2 chosen to determine their performance on the Czech retrieval dataset DaReCzech. The primary objective of our experiments is to estimate the quality of modern retrieval approaches in the Czech language. Our analyses include retrieval quality, speed, and memory footprint. Secondly, we analyze whether it is better to use the model directly in Czech text, or to use machine translation into English, followed by retrieval in English. Our experiments identify the most effective option for Czech information retrieval. The findings revealed notable performance differences among the models, with Gemma22 achieving the highest precision and recall, while Contriever performing poorly. Conclusively, SPLADE and PLAID models offered a balance of efficiency and performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。