arXiv:2509.00550cs.LGcs.CV2025-09被引 1

融合财务与文本数据的决策树模型,提升中小企业信贷评估准确率

Integrated Multivariate Segmentation Tree for Heterogeneous Credit Data Analysis in Small- and Medium-Sized Enterprises

  • 用矩阵分解将文本转为数值,结合Lasso选关键财务特征
  • 构建多变量决策树,88.9%准确率优于传统模型
  • 解释性强、计算快,适合需要可解释信贷评估的场景

传统决策树仅依赖数值变量,在处理高维数据和文本信息方面存在局限。为此,我们提出集成多变量分段树(IMST),通过三阶段框架提升中小企业(SMEs)信贷评估效果:首先利用矩阵分解将文本数据转化为数值矩阵;其次采用Lasso回归筛选关键财务特征;最后基于基尼系数或熵构建多变量决策树,并应用最弱链接剪枝控制模型复杂度。基于1428家中国中小企业数据的实验表明,IMST准确率达88.9%,高于基准决策树(87.4%)及支持向量机、神经网络等传统模型。该模型兼具更强可解释性与计算效率,结构更简洁,风险识别能力更优。

原文摘要 · Abstract (English)

Traditional decision tree models, which rely exclusively on numerical variables, often face challenges in handling high-dimensional data and are limited in their ability to incorporate textual information effectively. To address these limitations, we propose the integrated multivariate segmentation tree (IMST), a comprehensive framework designed to improve credit evaluation for small- and medium-sized enterprises (SMEs) by integrating financial data with textual sources. This method comprises three core stages: (1) transforming textual data into numerical matrices through matrix factorization, (2) selecting salient financial features using Lasso regression, and (3) constructing a multivariate segmentation tree based on either the Gini index or entropy, with weakest-link pruning applied to control model complexity. Experimental results based on a dataset of 1,428 Chinese SMEs demonstrated that IMST achieved an accuracy rate of 88.9%, surpassing both baseline decision trees (87.4%) and conventional models such as support vector machines and neural networks. Furthermore, the proposed model demonstrated superior interpretability and computational efficiency, featuring a more streamlined architecture and improved risk detection capabilities.

信贷评估决策树文本分析中小企业

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。