arXiv:2503.04644cs.CLcs.IR2025-03NAACL被引 19

首个专家领域指令跟随检索评测基准,覆盖四大专业场景

IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval

  • 构建包含2426个样本的多领域指令跟随检索数据集
  • 15个前沿模型在复杂指令下表现普遍不足,尤其在专业领域
  • 适合关注大模型信息检索能力评估的研究者与开发者

我们提出IFIR,首个针对专家领域指令跟随信息检索的综合性评测基准。IFIR包含2,426个高质量样本,覆盖金融、法律、医疗和科学文献四大专业领域,共八个子集。每个子集对应一种或多种领域特异性检索任务,模拟真实场景中定制化指令的关键作用。通过引入不同复杂度的指令,IFIR支持对指令跟随检索能力的细粒度分析。我们还提出一种基于大模型的新评估方法,以更精确、可靠地衡量模型遵循指令的表现。在15个前沿检索模型上的广泛实验表明,当前模型在处理复杂、领域特定指令时面临显著挑战。我们进一步提供了深入分析,揭示其局限性,为未来检索器的发展提供重要指导。

原文摘要 · Abstract (English)

We introduce IFIR, the first comprehensive benchmark designed to evaluate instruction-following information retrieval (IR) in expert domains. IFIR includes 2,426 high-quality examples and covers eight subsets across four specialized domains: finance, law, healthcare, and science literature. Each subset addresses one or more domain-specific retrieval tasks, replicating real-world scenarios where customized instructions are critical. IFIR enables a detailed analysis of instruction-following retrieval capabilities by incorporating instructions at different levels of complexity. We also propose a novel LLM-based evaluation method to provide a more precise and reliable assessment of model performance in following instructions. Through extensive experiments on 15 frontier retrieval models, including those based on LLMs, our results reveal that current models face significant challenges in effectively following complex, domain-specific instructions. We further provide in-depth analyses to highlight these limitations, offering valuable insights to guide future advancements in retriever development.

信息检索指令跟随评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。