arXiv:2604.13332cs.LG2026-04被引 1

用基础模型挖掘特征交互,提升可解释模型性能

Selecting Feature Interactions for Generalized Additive Models by Distilling Foundation Models

  • 用表格基础模型隐式学习特征依赖,再通过后处理提取重要交互
  • 在多个任务中,新发现的交互使GAM预测性能显著提升
  • 适合需要高可解释性且追求精度的表格数据建模场景

在表格数据建模中,识别有意义的特征交互是构建准确且可解释模型的核心挑战。广义加性模型(GAMs)虽在表格数据上表现优异,但通常依赖启发式方法选择交互项,可能遗漏高阶或上下文相关的效应。为此,我们提出TabDistill方法,利用表格基础模型和事后蒸馏技术。核心思路是:表格基础模型通过大规模表示学习隐式捕捉丰富的自适应特征依赖。给定数据集后,先拟合一个表格基础模型,再使用事后交互归因方法从中提取显著特征交互。随后将这些交互作为项用于GAM建模。在多个任务中,由TabDistill识别出的交互均带来下游GAM预测性能的稳定提升。结果表明,表格基础模型可作为高效的数据驱动引导工具,连接高容量模型与可解释的加性框架。

原文摘要 · Abstract (English)

Identifying meaningful feature interactions is a central challenge in building accurate and interpretable models for tabular data. Generalized additive models (GAMs) have shown great success at modeling tabular data, but often rely on heuristic procedures to select interactions, potentially missing higher-order or context-dependent effects. To meet this challenge, we propose TabDistill, a method that leverages tabular foundation models and post-hoc distillation methods. Our key intuition is that tabular foundation models implicitly learn rich, adaptive feature dependencies through large-scale representation learning. Given a dataset, TabDistill first fits a tabular foundation model to the dataset, and then applies a post-hoc interaction attribution method to extract salient feature interactions from it. We evaluate these interactions by then using them as terms in a GAM. Across tasks, we find that interactions identified by TabDistill lead to consistent improvements in downstream GAMs' predictive performance. Our results suggest that tabular foundation models can serve as effective, data-driven guides for interaction discovery, bridging high-capacity models and interpretable additive frameworks.

特征交互GAM基础模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。