构建首个支持可解释数据集搜索的评测基准
DSEBench: A Test Collection for Explainable Dataset Search with Examples
- 提出结合关键词与目标数据集的新型搜索任务DSE
- 建立首个带字段级标注的数据集搜索评测集
- 适合研究可解释信息检索与大模型应用的学者
数据集搜索是语义网与信息检索领域的经典任务。现有方法基于关键词或相似数据集检索,但在同时涉及关键词和目标数据集时表现不佳。为此,本文提出广义任务DSE(带示例的数据集搜索),并进一步发展为可解释的DSE(ExDSE),要求识别检索结果的相关领域。我们构建了DSEBench,首个提供高质量数据集级与字段级标注的测试集,分别支持DSE与ExDSE评估。利用大语言模型生成大量训练标注。在DSEBench上,我们通过适配和评估多种词法、稠密向量及大模型基检索、重排序与解释方法,建立了全面基线。
原文摘要 · Abstract (English)
Dataset search is a well-established task in the Semantic Web and information retrieval research. Current approaches retrieve datasets either based on keyword queries or by identifying datasets similar to a given target dataset. These paradigms fail when the information need involves both keywords and target datasets. To address this gap, we investigate a generalized task, Dataset Search with Examples (DSE), and extend it to Explainable DSE (ExDSE), which further requires identifying relevant fields of the retrieved datasets. We construct DSEBench, the first test collection that provides high-quality dataset-level and field-level annotations to support the evaluation of DSE and ExDSE, respectively. In addition, we employ a large language model to generate extensive annotations for training purposes. We establish comprehensive baselines on DSEBench by adapting and evaluating a variety of lexical, dense, and LLM-based retrieval, reranking, and explanation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。