arXiv:2510.21737cs.IR2025-10被引 1

构建首个跨表格与文本的数据产品发现基准,支持高阶分析需求。

DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text

  • 通过聚类表与文本生成数据产品,用大模型生成专业级请求。
  • 包含13,076个验证实例,覆盖开放域与金融领域。
  • 适合研究数据资产整合、智能分析系统开发的团队。

数据产品是为特定业务场景设计的可复用、自包含资源。自动化发现对工业界极具吸引力,能提升大规模数据湖中的数据访问效率并支撑分析工作流。然而,目前尚无针对混合表格-文本语料库的数据产品发现基准。现有数据集仅关注单个表格上的事实性问题回答,而非将多个相关数据资产整合为连贯的产品。为此,我们提出DPDisc,首个大规模数据产品发现基准,要求系统检索表与段落的连贯集合以满足高层级数据产品请求(DPRs)。我们引入DPForge自动化流水线,通过聚类相关表与段落形成数据产品,利用大模型集成生成专业级分析请求,并通过多阶段大模型评估验证质量。DPDisc包含13,076个经验证的实例,源自三个代表性的数据集,涵盖开放域与金融领域。基于稀疏、稠密及混合检索方法的基线实验表明评估可行性,同时揭示不同领域间显著性能差距,为结构感知的数据产品发现研究提供契机。代码与数据集见:数据集:https://huggingface.co/datasets/ibm-research/data-product-benchmark;代码:https://github.com/ibm/data-product-benchmark

原文摘要 · Abstract (English)

Data products are reusable, self-contained assets designed for specific business use cases. Automating their discovery is of great industry interest, as it enables efficient data access in large data lakes and supports analytical workflows. However, no benchmark currently exists for data product discovery over hybrid table-text corpora. Existing datasets focus on answering single factoid questions over individual tables rather than assembling multiple related data assets into coherent products. To address this gap, we present DPDisc, the first large-scale benchmark for data product discovery, where systems must retrieve coherent collections of tables and passages to satisfy high-level Data Product Requests (DPRs). We introduce DPForge, an automated pipeline that systematically repurposes table-text QA datasets by clustering related tables and passages into coherent data products, generating professional-level analytical requests using an LLM ensemble, and validating quality through multi-phase LLM evaluation. DPDisc comprises 13,076 validated instances with full provenance, derived from three representative datasets spanning open-domain and financial domains. Baseline experiments with sparse, dense, and hybrid retrieval methods imply evaluation feasibility while revealing substantial performance gaps across domains, indicating opportunities for future research in structure-aware data product discovery. Code and datasets are available at: Dataset: https://huggingface.co/datasets/ibm-research/data-product-benchmark Code: https://github.com/ibm/data-product-benchmark

数据发现大模型应用智能分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。