arXiv:2412.13102cs.IRcs.CL2024-12ACL被引 20

用大模型自动生成多领域多语言检索数据,解决评测滞后问题

AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark

  • 用大模型自动构建测试数据,无需人工标注
  • 覆盖多任务、多领域、多语言,数据多样性高
  • 动态更新,适合持续评估新兴检索模型

评估在信息检索(IR)模型发展中至关重要。然而,现有基准依赖预设领域和人工标注数据,在应对新兴领域时成本高、效率低。为此,我们提出自动化异构信息检索基准(AIR-Bench)。其三大特点:1)自动化——测试数据由大语言模型(LLMs)自动生成,无需人工参与;2)异构性——数据涵盖多样任务、领域和语言;3)动态性——持续扩展覆盖的领域与语言,为社区开发者提供日益全面的评测体系。我们构建了可靠的数据生成管道,基于真实语料自动生成多样化高质量评估数据集。实验表明,生成数据与人工标注数据高度一致,证明AIR-Bench可作为可信的IR模型评测基准。相关资源已公开:https://github.com/AIR-Bench/AIR-Bench。

原文摘要 · Abstract (English)

Evaluation plays a crucial role in the advancement of information retrieval (IR) models. However, current benchmarks, which are based on predefined domains and human-labeled data, face limitations in addressing evaluation needs for emerging domains both cost-effectively and efficiently. To address this challenge, we propose the Automated Heterogeneous Information Retrieval Benchmark (AIR-Bench). AIR-Bench is distinguished by three key features: 1) Automated. The testing data in AIR-Bench is automatically generated by large language models (LLMs) without human intervention. 2) Heterogeneous. The testing data in AIR-Bench is generated with respect to diverse tasks, domains and languages. 3) Dynamic. The domains and languages covered by AIR-Bench are constantly augmented to provide an increasingly comprehensive evaluation benchmark for community developers. We develop a reliable and robust data generation pipeline to automatically create diverse and high-quality evaluation datasets based on real-world corpora. Our findings demonstrate that the generated testing data in AIR-Bench aligns well with human-labeled testing data, making AIR-Bench a dependable benchmark for evaluating IR models. The resources in AIR-Bench are publicly available at https://github.com/AIR-Bench/AIR-Bench.

信息检索自动评测大模型应用多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。