评测大模型在数据仓库图结构推理能力,涵盖外键与数据血缘关系。
DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning

- 构建包含5个数据仓库模式的自动问答集,融合外键与数据血缘边。
- 工具增强方法显著优于静态方法,但在复杂组合任务上已达性能瓶颈。
- 适合研究数据库语义理解、大模型推理能力的学者参考。
本文提出DW-Bench,一个新基准,用于评估大语言模型(LLMs)在数据仓库模式上的图拓扑推理能力,明确整合了外键(FK)和数据血缘边。该基准包含1,046个自动生成且可验证正确的题目,覆盖五个不同的数据仓库模式。实验表明,工具增强的方法显著优于静态方法,但在困难的组合型子类型上表现趋于饱和。
原文摘要 · Abstract (English)
This paper introduces DW-Bench, a new benchmark that evaluates large language models (LLMs) on graph-topology reasoning over data warehouse schemas, explicitly integrating both foreign-key (FK) and data-lineage edges. The benchmark comprises 1,046 automatically generated, verifiably correct questions across five schemas. Experiments show that tool-augmented methods substantially outperform static approaches but plateau on hard compositional subtypes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。