构建126项任务的指令检索评测基准,评估大模型泛化能力
MAIR: A Massive Benchmark for Evaluating Instructed Retrieval
- 构建涵盖6个领域的126个异构检索任务,覆盖广泛场景
- 指令微调模型整体表现优于非指令微调模型,但在长尾任务上仍弱
- 适合评估大模型在指令驱动下的泛化与鲁棒性
近期信息检索(IR)模型通过大规模数据集和任务进行预训练与指令微调,能在多种任务上表现良好,并可能通过指令泛化到未见任务。然而,现有IR评测基准任务范围有限,难以全面评估最新模型。本文提出MAIR(Massive Instructed Retrieval Benchmark),一个包含6个领域共126个不同检索任务的异构评测基准,数据来自现有公开数据集。我们对当前最先进的指令微调文本嵌入模型和重排序模型进行了评测。实验表明,指令微调模型在MAIR上整体性能优于非指令微调模型。此外,结果提示当前指令微调的文本嵌入模型与重排序模型在特定长尾任务中仍缺乏有效性。MAIR已开源:https://github.com/sunnweiwei/Mair。
原文摘要 · Abstract (English)
Recent information retrieval (IR) models are pre-trained and instruction-tuned on massive datasets and tasks, enabling them to perform well on a wide range of tasks and potentially generalize to unseen tasks with instructions. However, existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models. In this paper, we propose MAIR (Massive Instructed Retrieval Benchmark), a heterogeneous IR benchmark that includes 126 distinct IR tasks across 6 domains, collected from existing datasets. We benchmark state-of-the-art instruction-tuned text embedding models and re-ranking models. Our experiments reveal that instruction-tuned models generally achieve superior performance compared to non-instruction-tuned models on MAIR. Additionally, our results suggest that current instruction-tuned text embedding models and re-ranking models still lack effectiveness in specific long-tail tasks. MAIR is publicly available at https://github.com/sunnweiwei/Mair.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。