arXiv:2608.26750cs.AI2026-08

用大模型破解数据湖中编码字段的关系发现难题

Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case

  • 分两阶段构建列嵌入,融合元数据与数据特征
  • 在工业级ERP数据上准确识别语义相关字段,优于传统方法
  • 适合处理标签模糊、信息稀疏的业务数据场景

数据湖依赖元数据维持可用性,但现有元数据常不足以支持列关系发现,尤其在包含编码或缩写命名的ERP数据中。本文提出ColRel,一种两阶段方法:第一阶段基于摄入时可得的元数据与数据构建列嵌入;第二阶段在困难情况下,利用业务词典帮助解析列名,并生成简短自然语言描述以增强理解。在公开基准和一个工业级ERP数据集上的实验表明,该方法在语义相关且信号弱的场景下表现优异。

原文摘要 · Abstract (English)

Data lakes rely on metadata to remain usable, yet this meta data is often limited or weakly informative for column relationship discovery, especially in ERP-derived datasets with coded or abbreviated schema labels. We propose ColRel, a two-stage method that builds column embeddings from metadata and data available at ingestion time. In difficult cases, such as coded schemata, business dictionaries help better interpret column names and support the generation of short natural-language descriptions used in the second stage. Experiments on public benchmarks and an industrial ERP dataset show that ColRel is particularly effective in semantically related, weak-signal settings.

数据湖大模型关系发现ERP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。