arXiv:2606.10333cs.LGcs.CR2026-06

在保护用户隐私前提下,安全融合手机等替代数据提升信贷风险预测精度。

Privacy-Preserving Credit Risk Prediction with Alternative Data

  • 提出PrivacyCredit方法,实现隐私保护下的跨机构数据协作建模。
  • 实验表明其预测性能与直接使用原始数据模型相当,误差仅0.3%。
  • 适合金融风控、数据安全合规领域研究人员和从业者参考。

信贷风险预测是消费金融行业的重要问题。传统上,金融机构依赖借款人的身份、财务及信用历史等传统数据构建预测模型。近期研究表明,手机通信等替代数据能更全面准确地刻画借款人信用状况,从而提升预测效果。然而,替代数据通常由金融机构外的第三方持有,直接共享会侵犯消费者隐私,而现有研究普遍忽视此问题。为此,我们提出隐私保护型替代数据信贷风险预测新范式,同时满足三项实际约束:隐私保护(保护用户隐私)、模型机密性(模型集中于金融机构学习存储)和无损性(保持模型性能)。我们开发了PrivacyCredit这一新型隐私保护机器学习方法,并从理论上证明其具备隐私保护、模型机密性和无损性。基于真实世界信贷数据集(含替代数据)的大量实验表明,安全融合替代数据可显著提升预测能力,且PrivacyCredit的预测性能与明文合并传统与替代数据训练的模型几乎一致(相对误差<0.3%)。此外,我们还验证了其模型机密性与计算效率。

原文摘要 · Abstract (English)

Credit risk prediction is a critical problem in the consumer credit industry. Traditionally, financial institutions construct credit risk prediction models using borrowers' demographic, financial, and credit history data, collectively referred to as traditional data. Recent studies have demonstrated that alternative data, such as borrowers' mobile phone communication data, enable lenders to acquire fuller and more accurate profiles of borrowers' creditworthiness, thereby improving credit risk prediction performance. Nevertheless, alternative data are held by external entities independent of financial institutions. Directly sharing alternative data with financial institutions infringe on consumer privacy, yet existing credit risk prediction studies largely overlook this issue. To address this gap, we define a new problem, namely privacy-preserving credit risk prediction with alternative data, which simultaneously considers three practical constraints: the privacy-preserving constraint that protects consumer privacy, the model-confidentiality constraint that learns and stores the model centrally at the financial institution, and the lossless constraint that maintains the performance of the learned model. To solve this problem, we develop PrivacyCredit, a novel privacy-preserving machine learning method. We then theoretically demonstrate the privacy-preserving, model-confidential, and lossless properties of PrivacyCredit. Through extensive experiments using a real-world credit dataset linked with alternative data, we demonstrate the predictive value of securely incorporating alternative data into credit risk prediction and show that PrivacyCredit achieves the same predictive performance as the model learned from the insecure plaintext combination of traditional and alternative data. We further evaluate its model-confidentiality property and computational efficiency.

信贷风险隐私计算替代数据联邦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。