arXiv:2511.10964cs.LGcs.AI2025-11

研究数据质量对信贷风险模型的影响,发现不同问题对模型表现影响差异大。

How Data Quality Affects Machine Learning Models for Credit Risk Assessment

  • 用可控污染方法测试10种模型在缺失值、噪声等数据问题下的表现
  • 部分模型在数据退化时准确率下降超30%,体现鲁棒性差异
  • 为从业者提供数据管道优化工具,适合关注数据可信性的研究者

机器学习模型在信贷风险评估中应用日益广泛,其效果高度依赖输入数据质量。本文研究缺失值、噪声属性、异常值和标签错误等数据质量问题对信贷风险预测模型性能的影响。基于开源数据集,利用Pucktrick库实施受控数据污染,评估了包括随机森林、SVM、逻辑回归在内的10种常用模型的鲁棒性。实验表明,模型对不同类型和严重程度的数据退化表现出显著差异的敏感性。所提出的框架与工具可为实践者提升数据流水线可靠性提供支持,并为数据驱动型AI研究者提供灵活的实验环境。

原文摘要 · Abstract (English)

Machine Learning (ML) models are being increasingly employed for credit risk evaluation, with their effectiveness largely hinging on the quality of the input data. In this paper we investigate the impact of several data quality issues, including missing values, noisy attributes, outliers, and label errors, on the predictive accuracy of the machine learning model used in credit risk assessment. Utilizing an open-source dataset, we introduce controlled data corruption using the Pucktrick library to assess the robustness of 10 frequently used models like Random Forest, SVM, and Logistic Regression and so on. Our experiments show significant differences in model robustness based on the nature and severity of the data degradation. Moreover, the proposed methodology and accompanying tools offer practical support for practitioners seeking to enhance data pipeline robustness, and provide researchers with a flexible framework for further experimentation in data-centric AI contexts.

信用风险数据质量机器学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。