arXiv:2605.22335cs.LG2026-05被引 1

让表格预测模型自动学习变量因果顺序,提升干预下的可靠性。

Learning Causal Orderings for In-Context Tabular Prediction

  • 用因果顺序约束注意力机制,只依赖前置变量做预测。
  • 在真实生物数据上恢复准确变量顺序,预测与缺失值填补效果佳。
  • 适合处理分布变化或干预场景的表格数据建模任务。

针对观测数据中的表格预测,现有方法主要依赖相关性,但在分布偏移或干预下表现不可靠。尽管已有因果结构发现方法,但通常与实际预测架构脱节。为此,我们提出TabOrder模型,通过无监督学习方式联合推断并强制变量的拓扑因果顺序。该模型使用因果顺序约束的注意力机制,仅基于目标变量之前的特征进行预测。类似于因果发现方法,其通过似然目标学习最优变量顺序。我们在标准函数模型类下证明了该设计的合理性,并研究了常见于表格数据的样本缺失如何影响因果方向识别。实验表明,TabOrder能准确恢复变量顺序,在预测与缺失值填补任务中表现良好,并为真实生物数据在干预下的分析提供洞见。

原文摘要 · Abstract (English)

In-context learning for tabular data sets strong predictive standards in observational settings; it however primarily relies on correlational structure, which becomes unreliable under distribution shift or intervention. While established methods to discover causal structure exist, they are often focused on structure identifiability and decoupled from the predictive architectures that could benefit from them. To bridge these perspectives, we study how to simultaneously infer and enforce causal structure in the form of topological variable orderings into tabular prediction. Unlike standard architectures, our model TabOrder uses causal order-constrained attention, basing predictions only on features that precede a target under a learned causal order. Similar to causal discovery methods, TabOrder learns the optimal variable ordering in an unsupervised manner through a likelihood-based objective. We justify this choice under standard functional model classes and also study how sample missingness, a common challenge in tabular data, interacts with causal direction identification. Empirically, we confirm that TabOrder recovers accurate variable orderings while addressing prediction and imputation tasks, as well as gives insight into real-world biological data under intervention.

因果学习表格数据预测建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。