arXiv:2605.18635cs.LGcs.AI2026-05

调教数据呈现方式,让表格大模型在信贷风险预测上超越传统方法。

Data Presentation Over Architecture: Resampling Strategies for Credit Risk Prediction with Tabular Foundation Models

论文配图:Data Presentation Over Architecture: Resampling Strategies for Credit Risk Prediction with Tabular Foundation Models
图 1 · 摘自论文原文
  • 通过平衡采样构建上下文,显著提升模型性能。
  • 5K-10K样本的均衡上下文使大模型达到经典模型效果。
  • 适合关注数据构建策略的风控算法工程师。

信贷违约预测是具有严重类别不平衡、特征异质性及严格延迟约束的表格学习问题。表格基础模型(TFMs)通过上下文学习应对该问题,其预测结果对上下文窗口的构建方式高度敏感。我们在Home Credit和Lending Club数据集上对比了四种经典模型与五种TFMs,测试七种上下文构建策略及1K至50K的上下文规模。结果显示,上下文策略对AUC-ROC的解释方差超过TFM架构选择:平衡与混合采样比均匀采样提升3至4个AUC点,且差距超过不同TFM家族间的差异。使用5K至10K样本的均衡上下文,最强的TFMs达到了全数据训练经典基线的AUC水平,并恢复了默认阈值GBDT忽略的重要违约类召回率。这表明在不平衡信贷风险场景中,上下文构建而非模型架构,是部署TFMs的关键杠杆。

原文摘要 · Abstract (English)

Credit default prediction is a tabular learning problem with severe class imbalance, heterogeneous features, and tight latency budgets. Tabular Foundation Models (TFMs) approach this problem through in-context learning, which makes their predictions sensitive to how the context window is built. We benchmark four classical models and five TFMs on the Home Credit and Lending Club datasets, varying the context-construction strategy (seven options) and the context size (1K to 50K). On both datasets, the choice of context strategy explains more variance in AUC-ROC than the choice of TFM family: balanced and hybrid sampling add 3 to 4 AUC points over uniform sampling, and the gap exceeds the spread between TFMs. With a balanced context of 5K to 10K examples, the strongest TFMs reach the AUC of classical baselines trained on the full data, while also recovering meaningful default-class recall that default-threshold GBDTs do not. We frame this as evidence that context construction, rather than architecture choice, is the primary deployment lever for TFMs in imbalanced credit-risk settings.

信贷风险表格模型上下文构建不平衡数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。