arXiv:2511.02593cs.LG2025-11被引 1

用大模型融合财务数据与机器学习,提升企业信用评分的准确性和可解释性。

A Large Language Model for Corporate Credit Scoring

  • 融合财务指标与语言模型,结合多种梯度提升算法优化预测。
  • 在7800家企业数据上测试,跨机构平均AUC超0.93,表现稳定可靠。
  • 适合金融风控、评级机构及需要透明决策依据的场景使用。

我们提出Omega^2,一种基于大语言模型的企业信用评分框架,通过整合结构化财务数据与先进机器学习方法,提升预测可靠性与可解释性。研究在包含7,800家企业的多机构数据集上评估,数据来自穆迪、标普、惠誉和埃根-琼斯,涵盖杠杆率、利润率、流动性比率等企业级财务指标。系统采用贝叶斯搜索优化的CatBoost、LightGBM和XGBoost模型,并在时间序列验证下确保结果具备前瞻性与可复现性。Omega^2在各机构测试中平均AUC超过0.93,证明其在不同评级体系间具有强泛化能力且时间一致性良好。结果表明,将语言推理与量化学习结合,可构建透明、符合机构标准的企业信用风险评估基础。

原文摘要 · Abstract (English)

We introduce Omega^2, a Large Language Model-driven framework for corporate credit scoring that combines structured financial data with advanced machine learning to improve predictive reliability and interpretability. Our study evaluates Omega^2 on a multi-agency dataset of 7,800 corporate credit ratings drawn from Moody's, Standard & Poor's, Fitch, and Egan-Jones, each containing detailed firm-level financial indicators such as leverage, profitability, and liquidity ratios. The system integrates CatBoost, LightGBM, and XGBoost models optimized through Bayesian search under temporal validation to ensure forward-looking and reproducible results. Omega^2 achieved a mean test AUC above 0.93 across agencies, confirming its ability to generalize across rating systems and maintain temporal consistency. These results show that combining language-based reasoning with quantitative learning creates a transparent and institution-grade foundation for reliable corporate credit-risk assessment.

信用评分大模型金融风控可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。