arXiv:2504.15849cs.IR2025-04中稿 · SIGIR'25被引 2

让自然语言条件与表格发现结合,提升精准检索能力

NLCTables: A Dataset for Marrying Natural Language Conditions with Table Discovery

  • 提出自然语言条件表发现新任务,用查询表+自然语言精炼结果
  • 构建含627个查询、2.2万张候选表的基准数据集
  • 适合做表格检索、人机交互或自然语言接口研究者使用

随着表格数据资源日益丰富,精准发现所需表格仍具挑战。现有方法多依赖查询表或模糊关键词,导致结果集过大需手动筛选。为此,我们提出新任务:自然语言条件表发现(nlcTD),允许用户结合查询表与自然语言要求来细化搜索。为此构建了nlcTables数据集,包含627个多样化查询(涵盖纯自然语言、并集、连接、模糊条件),22,080张候选表及21,200条相关性标注。在该数据集上对六种先进表格发现方法的评估显示显著性能差距,凸显该任务难度。数据集、构建框架及基线代码已公开于https://github.com/SuDIS-ZJU/nlcTables,以推动后续研究。

原文摘要 · Abstract (English)

With the growing abundance of repositories containing tabular data, discovering relevant tables for in-depth analysis remains a challenging task. Existing table discovery methods primarily retrieve desired tables based on a query table or several vague keywords, leaving users to manually filter large result sets. To address this limitation, we propose a new task: NL-conditional table discovery (nlcTD), where users combine a query table with natural language (NL) requirements to refine search results. To advance research in this area, we present nlcTables, a comprehensive benchmark dataset comprising 627 diverse queries spanning NL-only, union, join, and fuzzy conditions, 22,080 candidate tables, and 21,200 relevance annotations. Our evaluation of six state-of-the-art table discovery methods on nlcTables reveals substantial performance gaps, highlighting the need for advanced techniques to tackle this challenging nlcTD scenario. The dataset, construction framework, and baseline implementations are publicly available at https://github.com/SuDIS-ZJU/nlcTables to foster future research.

表格发现自然语言数据检索基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。