揭秘表格类基础模型的内在机制与失效原因
A Mechanistic Study of Tabular Foundation Models
- 不同架构模型用相似性读出策略,从注意力投票到类别均值
- 移除特定位置参数可使排列不变性精确化且不损失精度
- 针对读出机制设计攻击能复现预测失败,揭示模型弱点
具有不同架构的表格基础模型在多种分类和回归任务中收敛于相近的准确率,引发深层问题:(i) 模型是否执行相同的上下文算法?(ii) 行、列、类别排列不变性来自何处?(iii) 在针对推断机制设计的扰动下其鲁棒性如何?我们系统解答了这三个问题。模型家族展现出本质不同的基于相似性的读出机制:从上下文标签的注意力加权投票到类别条件均值读出,均通过因果干预验证。我们发现先前研究强调的表征坍缩对这些模型并非实际问题。每种模型的排列不变性源自特定的位置参数,移除后仍保持精度,且使近似不变性变为精确。针对每种读出机制设计的扰动成功复现预测的失败模式;枢纽攻击与秩攻击将它们与可重训练基线区分开。这些结果为当代表格基础模型提供了机制解释,并识别出决定其准确性和典型失败的关键归纳偏置。
原文摘要 · Abstract (English)
Tabular foundation models with different architectures converge in accuracy across a range of classification and regression tasks. This raises questions a leaderboard cannot answer: (i) whether the models execute the same in-context algorithm, (ii) where row, column, and class-permutation invariances originate, and (iii) how robust they are under perturbations engineered against the inferred mechanism. We characterize all three. The model families realize qualitatively distinct similarity-based readouts: from an attention-weighted vote over context labels to a class-conditional mean readout, each confirmed by causal intervention. We find that the representation collapse highlighted in prior work is not a practical concern for them. Each model's permutation invariances trace to specific positional parameters whose removal preserves accuracy and makes approximate invariance exact. Perturbations engineered against each readout reproduce predicted failure modes; hub and rank attacks isolate them from refit baselines. Together these results give a mechanistic account of contemporary tabular foundation models and identify which inductive biases govern both their accuracy and characteristic failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。