构建跨类别检索新基准,用大模型降低评测成本
MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval

- 基于真实用户查询和大模型生成,构建多类别检索数据集
- 覆盖文献、数据、变量、工具四类,支持统一评估
- 适合研究跨域检索与评测的学者使用
用户日益期望现代搜索系统能通过统一接口无缝获取多元数据源与格式的信息。然而,当前信息检索评测基准未能跟上这一趋势,主要因缺乏反映当代搜索领域多样性的测试集。为此,我们提出MIRA,一个基于大规模社会科学搜索平台的新基准。MIRA旨在单一统一框架下实现对四类异构内容——文献、研究数据、变量、仪器与工具——的类别感知排序。该数据集具有三大特点:(1) 基于真实用户查询,评估更贴近实际;(2) 覆盖四个不同学术类别,支持多维度评测;(3) 利用大语言模型生成主题描述、叙事内容及相关性判断,显著降低测试集构建的人力与成本。本研究已发布该资源,为多面性、类别感知、集成式或跨类别信息检索研究提供基础评测平台。
原文摘要 · Abstract (English)
Users increasingly expect modern search systems to offer a unified interface that seamlessly retrieves information from diverse data sources and formats. However, current information retrieval (IR) evaluation benchmarks have not kept pace with this development, primarily due to the lack of test collections that represent the diversity of contemporary search domains. We address this critical gap with MIRA, a novel benchmark based on a large-scale social science search platform. MIRA is designed for category-aware ranking across heterogeneous categories - Publications, Research Data, Variables, and Instruments & Tools - within a single, unified evaluation framework. The proposed collection is distinctive in several ways: (1) it is built upon real user queries, providing a more realistic basis for evaluation; (2) it covers scholarly items from four distinct categories, enabling multi-faceted evaluation; and (3) it leverages a Large Language Model to generate topic descriptions and narratives, as well as for relevance assessment with respect to these topics, substantially reducing the labor and cost of test collection generation. We release this resource to benefit the community by providing a foundational testbed for the research on multi-faceted, category-aware, integrated, or cross-category information retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。