首个自动生成数据产品的基准测试,连接业务需求与结构化数据。
Bridging Business Intent and Data: A Benchmark for Automatic Relational Data Product Generation
- 构建融合ELT与文本转SQL数据集的DP-Bench基准
- 用大模型实现从自然语言到数据产品的端到端生成
- 适合数据工程与AI交叉研究者参考
数据产品通过将原始数据转化为可操作的资产来满足特定业务需求。尽管文本转SQL和ELT流水线已有进展,但尚无全面基准用于评估从高层业务请求自动生成数据产品的全过程。为此,我们提出DP-Bench,首个面向自动数据产品创建的基准,由现有ELT和文本转SQL数据集元素组合而成。同时,我们提出基于大语言模型的基线方法,为连接自然语言业务意图与结构化数据表示提供研究基础。相关数据集与代码已开源:https://huggingface.co/datasets/ibm-research/dp-bench。
原文摘要 · Abstract (English)
A data product is designed to address a specific business need by transforming raw data into a curated, usable asset that delivers actionable insights. Despite practical advances in related areas like text-to-SQL and ELT pipelines, there is no comprehensive benchmark for evaluating the end-to-end process of automatically generating such data products from high-level business requests. To fill this gap, we introduce DP-Bench, a first-of-its-kind benchmark for automatic data product creation, built by combining elements from existing ELT and text-to-SQL datasets. We also propose baseline methods using LLMs to generate data products, providing a foundation for future research in bridging natural language business intent and structured data representation. We make the DP-Bench dataset and code available at: https://huggingface.co/datasets/ibm-research/dp-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。