构建企业表格语义理解基准,推动模型融合上下文知识进行推理。
SALT-KG: A Benchmark for Semantics-Aware Learning on Enterprise Tables
- 将企业表格与元数据知识图谱关联,实现表间语义联动。
- 引入结构化业务知识后,模型在关系推理中仍存在明显语义理解短板。
- 适合研究大模型在企业级结构化数据中融合领域知识的学者使用。
基于面向关系预测的SALT基准(Klein等,2024),我们提出SALT-KG,一个面向企业表格语义感知学习的基准。SALT-KG通过将多表交易数据与以元数据知识图谱(OBKG)形式表示的结构化运营业务知识相链接,捕捉字段级描述、关系依赖和业务对象类型。这一扩展使模型能够联合推理表格证据与上下文语义,这对基础模型处理结构化数据日益关键。实证分析表明,尽管元数据特征在传统预测指标上仅带来小幅提升,但它们持续揭示了模型在利用关系上下文语义方面的缺陷。通过将表格预测重新定义为语义条件推理,SALT-KG建立了一个基准,推动基于声明性知识的表格基础模型发展,为大规模企业级结构化数据的语义关联提供了首个实证路径。
原文摘要 · Abstract (English)
Building upon the SALT benchmark for relational prediction (Klein et al., 2024), we introduce SALT-KG, a benchmark for semantics-aware learning on enterprise tables. SALT-KG extends SALT by linking its multi-table transactional data with a structured Operational Business Knowledge represented in a Metadata Knowledge Graph (OBKG) that captures field-level descriptions, relational dependencies, and business object types. This extension enables evaluation of models that jointly reason over tabular evidence and contextual semantics, an increasingly critical capability for foundation models on structured data. Empirical analysis reveals that while metadata-derived features yield modest improvements in classical prediction metrics, these metadata features consistently highlight gaps in the ability of models to leverage semantics in relational context. By reframing tabular prediction as semantics-conditioned reasoning, SALT-KG establishes a benchmark to advance tabular foundation models grounded in declarative knowledge, providing the first empirical step toward semantically linked tables in structured data at enterprise scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。