更长数据周期和更多特征反而降低房贷违约预测效果
Time Series Feature Redundancy Paradox: An Empirical Study Based on Mortgage Default Prediction
- 用短周期+精选特征提升预测精度
- 2012-2022长周期数据使准确率下降12.3%
- 适合金融风控中特征与时间窗优化
机器学习在金融风控中广泛应用,传统观点认为更长训练周期和更多特征变量能提升模型性能。本文以房利美(Fannie Mae)房贷数据为基础,实证发现时间序列预测中这一认知存在悖论:延长训练数据时间跨度及增加非关键特征反而显著降低预测效果。通过对比不同时间窗口(如2012–2022年)和特征组合的预测表现,研究发现单年周期搭配精心筛选的关键特征能获得更优结果。实验表明,过长的时间跨度引入了历史噪声与过时市场模式,而过多非关键特征干扰模型对核心违约因素的学习。该研究挑战了数据建模中‘越多越好’的传统观念,为金融风险预测中的特征选择与时间窗口优化提供了新思路与实践指导。
原文摘要 · Abstract (English)
With the widespread application of machine learning in financial risk management, conventional wisdom suggests that longer training periods and more feature variables contribute to improved model performance. This paper, focusing on mortgage default prediction, empirically discovers a phenomenon that contradicts traditional knowledge: in time series prediction, increased training data timespan and additional non-critical features actually lead to significant deterioration in prediction effectiveness. Using Fannie Mae's mortgage data, the study compares predictive performance across different time window lengths (2012-2022) and feature combinations, revealing that shorter time windows (such as single-year periods) paired with carefully selected key features yield superior prediction results. The experimental results indicate that extended time spans may introduce noise from historical data and outdated market patterns, while excessive non-critical features interfere with the model's learning of core default factors. This research not only challenges the traditional "more is better" approach in data modeling but also provides new insights and practical guidance for feature selection and time window optimization in financial risk prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。