arXiv:2602.00120cs.LG2026-02

用AutoML预测房贷违约,解决标签模糊、数据失衡和时间泄露问题。

Predicting Mortgage Default with Machine Learning: AutoML, Class Imbalance, and Leakage Control

  • 设计防泄漏特征选择与严格时序划分,确保模型评估真实可靠。
  • 在不同正负样本比下,模型性能稳定,AutoGluon表现最佳(AUROC最高)。
  • 适合金融风控从业者和机器学习应用研究者参考实操细节。

房贷违约预测是金融风险管理的核心任务,机器学习模型被广泛用于估算违约概率并提供可解释信号。然而,在真实贷款数据中,三个因素常导致评估无效与部署不可靠:违约标签定义模糊、严重类别不平衡,以及由时间结构和事件后变量引发的信息泄露。本文基于真实贷款级数据集,比较多种机器学习方法,重点控制信息泄露与类别不平衡。采用防泄露特征选择、严格时序划分(限制放款与报告期)、以及受控的多数类下采样策略。在多个正负样本比例下,模型性能保持稳定,其中自动化机器学习方法(AutoGluon)在所有评估模型中取得最优的AUROC。本工作的扩展与教学版本将作为书籍章节发表。

原文摘要 · Abstract (English)

Mortgage default prediction is a core task in financial risk management, and machine learning models are increasingly used to estimate default probabilities and provide interpretable signals for downstream decisions. In real-world mortgage datasets, however, three factors frequently undermine evaluation validity and deployment reliability: ambiguity in default labeling, severe class imbalance, and information leakage arising from temporal structure and post-event variables. We compare multiple machine learning approaches for mortgage default prediction using a real-world loan-level dataset, with emphasis on leakage control and imbalance handling. We employ leakage-aware feature selection, a strict temporal split that constrains both origination and reporting periods, and controlled downsampling of the majority class. Across multiple positive-to-negative ratios, performance remains stable, and an AutoML approach (AutoGluon) achieves the strongest AUROC among the models evaluated. An extended and pedagogical version of this work will appear as a book chapter.

房贷违约AutoML数据泄露不平衡分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。