arXiv:2412.20331cs.DBcs.AI2024-12被引 18

LLM在企业数据上表现差,新基准和三招提升性能。

Mind the Data Gap: Bridging LLMs to Enterprise Data Integration

  • 用分层标注、运行时学习与本体合成提升LLM企业数据能力
  • 部署后企业数据性能达到公开数据水平
  • 适合关注企业级AI落地的研究者与工程师

主流大语言模型基于公开数据训练,但全球多数数据为不可见的暗数据,主要以私有组织或企业数据形式存在。我们发现,当在真实企业数据集上测试时,基于LLM的方法性能显著下降;现有基于公开数据的基准高估了其实际表现。为此,我们发布新基准Goby Benchmark,推动企业数据集成研究。基于该基准经验,提出三项技术:(1) 分层标注,(2) 运行时类别学习,(3) 本体合成。实验表明,部署这些技术后,企业数据上的性能可达到公开数据水平。Goby基准可通过 https://goby-benchmark.github.io/ 获取。

原文摘要 · Abstract (English)

Leading large language models (LLMs) are trained on public data. However, most of the world's data is dark data that is not publicly accessible, mainly in the form of private organizational or enterprise data. We show that the performance of methods based on LLMs seriously degrades when tested on real-world enterprise datasets. Current benchmarks, based on public data, overestimate the performance of LLMs. We release a new benchmark dataset, the GOBY Benchmark, to advance discovery in enterprise data integration. Based on our experience with this enterprise benchmark, we propose techniques to uplift the performance of LLMs on enterprise data, including (1) hierarchical annotation, (2) runtime class-learning, and (3) ontology synthesis. We show that, once these techniques are deployed, the performance on enterprise data becomes on par with that of public data. The Goby benchmark can be obtained at https://goby-benchmark.github.io/.

企业数据LLM基准测试数据集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。