arXiv:2504.12879cs.IRcs.CL2025-04被引 4

构建首个俄语信息检索评估基准,推动中文之外的多语言研究

Building Russian Benchmark for Evaluation of Information Retrieval Models

  • 整合17个俄语数据集,支持零样本评估
  • mE5-large等神经模型表现优于传统方法
  • 适合关注俄语、多语言检索的研究者

我们提出RusBEIR,一个面向俄语信息检索(IR)模型的零样本评估综合基准。该基准包含17个来自不同领域的数据集,涵盖已适配、翻译和新创建的数据,支持词汇模型与神经模型的系统性对比。研究表明,形态丰富的语言中预处理对词汇模型至关重要,且BM25在全文检索中仍是强基线。mE5-large和BGE-M3等神经模型在多数数据集上表现更优,但受输入长度限制,在长文档检索中面临挑战。RusBEIR提供统一、开源框架,推动俄语信息检索研究发展。

原文摘要 · Abstract (English)

We introduce RusBEIR, a comprehensive benchmark designed for zero-shot evaluation of information retrieval (IR) models in the Russian language. Comprising 17 datasets from various domains, it integrates adapted, translated, and newly created datasets, enabling systematic comparison of lexical and neural models. Our study highlights the importance of preprocessing for lexical models in morphologically rich languages and confirms BM25 as a strong baseline for full-document retrieval. Neural models, such as mE5-large and BGE-M3, demonstrate superior performance on most datasets, but face challenges with long-document retrieval due to input size constraints. RusBEIR offers a unified, open-source framework that promotes research in Russian-language information retrieval.

信息检索俄语基准测试神经模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。