arXiv:2601.00428cs.LG2026-01

对比16种可解释模型,发现回归任务表现有规律,分类却难预测。

Interpretable ML Under the Microscope: Performance, Meta-Features, and the Regression-Classification Predictability Gap

  • 用216个真实数据集评估16种可解释模型,分析性能与数据特征关系。
  • 回归任务中EBM和符号回归表现稳定,分类任务无可靠性能排序。
  • 显式追求简洁结构的模型训练慢得多,存在可解释性代价。

随着机器学习在高风险领域广泛应用,可解释性需求日益迫切。然而,针对表格数据的内在可解释模型的系统评估仍稀缺,且多仅关注整体性能。为此,我们在216个真实表格数据集上评估了16种可解释方法,包括解释性提升机(EBM)、符号回归(SR)和广义最优稀疏决策树,考察其预测准确率、计算效率及分布偏移下的泛化能力。超越平均性能排名,我们进一步分析模型表现如何随数据集元特征变化,并将其用于算法选择。结果揭示显著分化:回归任务中模型性能呈现可预测的层级,主要由EBM和SR主导,可由数据特征推断;而分类任务性能高度依赖具体数据集,无稳定层级,标准复杂度指标无法提供有效指导。此外,我们识别出“可解释性税”现象——显式优化结构稀疏性的模型训练时间显著更长。这些发现为寻求可解释性与预测性能平衡的实践者提供实用建议,并深化了对表格数据可解释建模的实证理解。

原文摘要 · Abstract (English)

As machine learning models are increasingly deployed in high-stakes domains, the need for interpretability has grown to meet strict regulatory and accountability constraints. Despite this interest, systematic evaluations of inherently interpretable models for tabular data remain scarce and often focus solely on aggregated performance. To address this gap, we evaluate sixteen interpretable methods, including Explainable Boosting Machines (EBMs), Symbolic Regression (SR), and Generalized Optimal Sparse Decision Trees, across 216 real-world tabular datasets. We assess predictive accuracy, computational efficiency, and generalization under distributional shifts. Moving beyond aggregate performance rankings, we further analyze how model behavior varies with dataset meta-features and operationalize these descriptors to study algorithm selection. Our analyses reveal a clear dichotomy: in regression tasks, models exhibit a predictable performance hierarchy dominated by EBMs and SR that can be inferred from dataset characteristics. In contrast, classification performance remains highly dataset-dependent with no stable hierarchy, showing that standard complexity measures fail to provide actionable guidance. Furthermore, we identify an "interpretability tax", showing that models explicitly optimizing for structural sparsity incur significantly longer training times. Overall, these findings provide practical guidance for practitioners seeking a balance between interpretability and predictive performance, and contribute to a deeper empirical understanding of interpretable modeling for tabular data.

可解释性表格数据模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。