企业表格数据与公开基准差异大,现有模型表现不一致。
Exploring Differences Between Tabular Enterprise Data and Public Benchmarks
- 对比企业数据与公开基准的统计特性差异
- 发现主流模型在真实企业数据上表现可能大幅下降
- 呼吁建立具备企业级特征的新基准
表格数据主导数据科学领域,吸引众多机器学习模型与专用基准。然而,对企业数据——业务运营的核心——了解甚少。为拓展商业应用的基准体系,本文分析了企业数据的统计特征及若干表格模型(如TabPFN、TabICL和ConTextTab)的性能表现。研究发现,企业数据与典型表格基准存在显著差异,且在公开基准上表现优异的模型,在真实企业数据上可能表现不佳,反之亦然。这一泛化能力的缺失凸显了构建具备企业级特性的新基准的必要性。
原文摘要 · Abstract (English)
Tabular data dominate the landscape of data science, increasingly attracting innovative machine learning models and tailored benchmarks. Yet, little is known for enterprise data, where tables constitute the backbone of business operations. To broaden the benchmarking landscape for business applications, this work aims to actualize the characteristics of enterprise data by providing an analysis of data statistics and performance measurements of tabular models such as TabPFN, TabICL and ConTextTab. Through our analysis, we find enterprise data markedly differ from tabular benchmarks and we demonstrate that a tabular model that performs well on typical tabular benchmarks may perform poorly on real world enterprise data -- and vice versa. This lack of generalization underlines the need for additional benchmarks with enterprise-grade characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。