统一检索评估框架,让研究者一键完成标准评测。
SuiteEval: Simplifying Retrieval Benchmarks
- 自动端到端评测,只需提供模型管道生成器
- 支持5大主流数据集,动态索引减少磁盘占用
- 适合需要可复现评测的检索与大模型研究者
信息检索评估常因数据集子集、聚合方式和流水线配置不一而难以复现与比较,尤其对需强跨域性能的基础嵌入模型而言。我们提出SuiteEval,一个统一框架,支持自动端到端评估、动态索引(复用磁盘索引以最小化存储)及主流基准(BEIR、LoTTE、MS MARCO、NanoBEIR、BRIGHT)的内置支持。用户仅需提供流水线生成器,框架即自动完成数据加载、索引构建、排序、指标计算与结果聚合。新基准套件可单行添加。SuiteEval降低重复工作,标准化评估流程,推动可复现的检索研究,契合日益扩展的基准需求。
原文摘要 · Abstract (English)
Information retrieval evaluation often suffers from fragmented practices -- varying dataset subsets, aggregation methods, and pipeline configurations -- that undermine reproducibility and comparability, especially for foundation embedding models requiring robust out-of-domain performance. We introduce SuiteEval, a unified framework that offers automatic end-to-end evaluation, dynamic indexing that reuses on-disk indices to minimise disk usage, and built-in support for major benchmarks (BEIR, LoTTE, MS MARCO, NanoBEIR, and BRIGHT). Users only need to supply a pipeline generator. SuiteEval handles data loading, indexing, ranking, metric computation, and result aggregation. New benchmark suites can be added in a single line. SuiteEval reduces boilerplate and standardises evaluations to facilitate reproducible IR research, as a broader benchmark set is increasingly required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。