揭秘表格基础模型如何内部处理数据,发现其隐藏表征可解释。
TabPFN Through The Looking Glass: An interpretability study of TabPFN and its internal representations
- 通过探测实验分析模型各层隐含表示中的信息
- 早期层已编码线性回归系数和中间计算结果
- 揭示模型决策过程的可解释性,适合研究者参考
表格基础模型是为多种表格数据任务设计的预训练模型,在多个领域表现出色,但其内部表征与学习到的概念仍不清晰。这种可解释性缺失使得研究模型如何处理和转换输入特征变得尤为重要。本文分析模型隐藏表征中编码的信息,并考察这些表征在不同层间的演化过程。我们开展了一系列探测实验,检验早期层中是否存在线性回归系数、复杂表达式的中间值以及最终答案。实验结果表明,表格基础模型的表征中存储了有意义且结构化的信息,能清晰识别出预测过程中涉及的中间量与最终输出。这为理解模型如何逐步优化输入及生成最终输出提供了洞见。研究结果深化了对表格基础模型内部机制的理解,证明其能编码具体可解释的信息,推动决策过程向更透明可信的方向发展。
原文摘要 · Abstract (English)
Tabular foundational models are pre-trained models designed for a wide range of tabular data tasks. They have shown strong performance across domains, yet their internal representations and learned concepts remain poorly understood. This lack of interpretability makes it important to study how these models process and transform input features. In this work, we analyze the information encoded inside the model's hidden representations and examine how these representations evolve across layers. We run a set of probing experiments that test for the presence of linear regression coefficients, intermediate values from complex expressions, and the final answer in early layers. These experiments allow us to reason about the computations the model performs internally. Our results provide evidence that meaningful and structured information is stored inside the representations of tabular foundational models. We observe clear signals that correspond to both intermediate and final quantities involved in the model's prediction process. This gives insight into how the model refines its inputs and how the final output emerges. Our findings contribute to a deeper understanding of the internal mechanics of tabular foundational models. They show that these models encode concrete and interpretable information, which moves us closer to making their decision processes more transparent and trustworthy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。