arXiv:2508.10053cs.LGstat.ML2025-08被引 15

xRFM融合核方法与树结构,提升表格数据建模精度与可解释性。

xRFM: Accurate, scalable, and interpretable feature learning models for tabular data

  • 结合特征学习核机器与树结构,自适应数据局部特征。
  • 在100个回归数据集上优于31种方法,200个分类任务中表现顶尖。
  • 天然提供可解释性,适合需透明决策的工业场景。

表格数据(由连续与类别变量构成的矩阵)的推断是现代科技与科学的基础。然而,与人工智能其他领域的迅猛发展相比,此类预测任务的最佳实践仍主要依赖梯度提升决策树(GBDTs)的变体,进展缓慢。近期,基于神经网络和特征学习的新方法重新激发了对表格数据先进方法的研究兴趣。本文提出xRFM,该算法将特征学习核机器与树结构结合,既能适应数据的局部结构,又可扩展至近乎无限规模的训练数据。实验表明,在100个回归数据集上,xRFM优于31种现有方法(包括新提出的表格式基础模型TabPFNv2和GBDTs),在200个分类数据集上表现竞争力,并超越传统GBDTs。此外,xRFM通过平均梯度外积(Average Gradient Outer Product)原生支持可解释性。

原文摘要 · Abstract (English)

Inference from tabular data, collections of continuous and categorical variables organized into matrices, is a foundation for modern technology and science. Yet, in contrast to the explosive changes in the rest of AI, the best practice for these predictive tasks has been relatively unchanged and is still primarily based on variations of Gradient Boosted Decision Trees (GBDTs). Very recently, there has been renewed interest in developing state-of-the-art methods for tabular data based on recent developments in neural networks and feature learning methods. In this work, we introduce xRFM, an algorithm that combines feature learning kernel machines with a tree structure to both adapt to the local structure of the data and scale to essentially unlimited amounts of training data. We show that compared to $31$ other methods, including recently introduced tabular foundation models (TabPFNv2) and GBDTs, xRFM achieves best performance across $100$ regression datasets and is competitive to the best methods across $200$ classification datasets outperforming GBDTs. Additionally, xRFM provides interpretability natively through the Average Gradient Outer Product.

表格数据可解释性特征学习核方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。