为表格模型提供可验证的特征重要性检验方法
Valid Feature-Level Inference for Tabular Foundation Models via the Conditional Randomization Test
- 结合条件随机化检验与TabPFN,实现特征层面的统计推断
- 在非线性相关场景下仍能给出有限样本有效的p值
- 无需重训练模型或假设分布,适合数据科学家快速验证
现代机器学习模型表达能力强但难以进行统计分析。尽管黑箱预测器表现优异,却很少提供有效假设检验或p值来判断单个特征是否包含目标变量信息。本文提出一种实用方法,将条件随机化检验(CRT)与TabPFN——一种用于表格数据的概率基础模型相结合。该方法在非线性与特征相关场景下,无需模型重训练或参数假设,即可获得条件特征相关性的有限样本有效p值,实现特征级别的可靠统计推断。
原文摘要 · Abstract (English)
Modern machine learning models are highly expressive but notoriously difficult to analyze statistically. In particular, while black-box predictors can achieve strong empirical performance, they rarely provide valid hypothesis tests or p-values for assessing whether individual features contain information about a target variable. This article presents a practical approach to feature-level hypothesis testing that combines the Conditional Randomization Test (CRT) with TabPFN, a probabilistic foundation model for tabular data. The resulting procedure yields finite-sample valid p-values for conditional feature relevance, even in nonlinear and correlated settings, without requiring model retraining or parametric assumptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。