arXiv:2506.21182cs.CLcs.AI2025-06被引 9

保障文本嵌入基准长期可用性与可复现性

Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks

  • 构建持续集成流水线确保数据完整性和测试自动化
  • 通过设计优化提升基准的可复现性与易用性
  • 支持社区贡献,持续扩展新任务与数据集

大规模文本嵌入基准(MTEB)已成为文本嵌入模型的标准评估平台。尽管先前工作已确立核心基准方法,本文聚焦于保障MTEB长期可复现性与可扩展性的工程实践。我们提出一套稳健的持续集成流程,实现数据完整性验证、测试自动化执行及评估结果泛化性分析。详细阐述了多项设计选择,共同提升可复现性与可用性。同时讨论了处理社区贡献及新增任务与数据集的策略。这些工程实践推动MTEB不断扩展,保持高质量与领域相关性。我们的经验为机器学习评估框架的维护者提供了宝贵参考。MTEB仓库地址:https://github.com/embeddings-benchmark/mteb

原文摘要 · Abstract (English)

The Massive Text Embedding Benchmark (MTEB) has become a standard evaluation platform for text embedding models. While previous work has established the core benchmark methodology, this paper focuses on the engineering aspects that ensure MTEB's continued reproducibility and extensibility. We present our approach to maintaining robust continuous integration pipelines that validate dataset integrity, automate test execution, and assess benchmark results' generalizability. We detail the design choices that collectively enhance reproducibility and usability. Furthermore, we discuss our strategies for handling community contributions and extending the benchmark with new tasks and datasets. These engineering practices have been instrumental in scaling MTEB to become more comprehensive while maintaining quality and, ultimately, relevance to the field. Our experiences offer valuable insights for benchmark maintainers facing similar challenges in ensuring reproducibility and usability in machine learning evaluation frameworks. The MTEB repository is available at: https://github.com/embeddings-benchmark/mteb

基准测试可复现性工程实践

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。