让表格大模型学会应对数据策略性操纵,提升真实场景预测准确率。
When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

- 推理时构建策略性上下文样本,动态对齐预测分布
- 在真实与合成数据上显著降低策略操纵带来的偏差
- 无需重训练,适合金融信贷等需应对人为干预的场景
基于预训练先验数据拟合网络(PFN)的表格基础模型在多种表格任务中表现出强泛化能力,但通常假设数据分布独立于部署分类器。在现实决策场景中,个体可能在模型部署后策略性修改自身特征以获得有利结果,导致部署后分布偏移。本文研究了此类策略性表格数据下PFN模型的泛化能力。结果表明,策略性操纵使预训练阶段学习到的非策略先验与部署后的策略先验不匹配,引发系统性预测偏差。为此,我们提出策略性先验数据拟合网络(SPN),一种无需重训练的推理时策略感知框架。SPN通过构造策略性上下文样本近似部署后输入,并对齐PFN预测与诱导出的策略分布。在真实世界与合成表格数据集上的实验表明,相比传统表格基础模型与经典方法,SPN在策略操纵下持续提升鲁棒性与预测性能。
原文摘要 · Abstract (English)
Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are typically designed for \emph{non-strategic} settings where data distributions are independent of deployed classifiers. In many real-world decision scenarios, however, individuals may strategically modify their features after deployment to obtain favorable outcomes, inducing a post-deployment distribution shift. This paper studies whether PFN-style tabular foundation models can generalize to such \emph{strategic} tabular data. We show that strategic manipulation creates a mismatch between the non-strategic prior learned during pretraining and the post-manipulation strategic prior, which leads to systematic prediction bias. To address this issue, we propose \textbf{Strategic Prior-data Fitted Network}~\textit{(SPN)}, an inference-time strategy-aware framework that adapts tabular foundation models to strategic environments without retraining. SPN constructs strategic in-context examples to approximate post-manipulation inputs and aligns PFN predictions with the induced strategic distribution. Experiments on real-world and synthetic tabular datasets show that SPN consistently improves robustness and predictive performance under strategic manipulation compared with both tabular foundation models and classical tabular methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。