融合财务与文本数据的聚类方法,提升信用风险识别精度
Advanced spectral clustering for heterogeneous data in credit risk monitoring systems
- 用优化权重融合财务与文本相似性,新方法选特征向量
- 聚类效果比单一数据方法高18%的轮廓系数,低风险企业30%未违约
- 适用于中小企业信用监控,能发现如'社会招聘'等关键线索
异构数据包含数值型财务变量和文本记录,对信用监控构成挑战。为此,我们提出先进谱聚类(ASC),通过优化权重整合财务与文本相似性,并采用新型特征值-轮廓系数优化方法选择特征向量。在包含1,428家中小企业的数据集上,ASC的轮廓系数比单类型数据基线方法高出18%。结果集群提供可操作洞察:51%的低风险企业文本中包含‘social recruitment’。ASC在k-means、k-medians、k-medoids等多种算法下均表现稳健,ΔIntra/Inter < 0.13,ΔSilhouette Coefficient < 0.02。该方法将谱聚类理论与异构数据应用结合,识别出如以招聘为核心业务的中小企业,其违约风险低30%,支持更精准有效的信用干预。
原文摘要 · Abstract (English)
Heterogeneous data, which encompass both numerical financial variables and textual records, present substantial challenges for credit monitoring. To address this issue, we propose Advanced Spectral Clustering (ASC), a method that integrates financial and textual similarities through an optimized weight parameter and selects eigenvectors using a novel eigenvalue-silhouette optimization approach. Evaluated on a dataset comprising 1,428 small and medium-sized enterprises (SMEs), ASC achieves a Silhouette score that is 18% higher than that of a single-type data baseline method. Furthermore, the resulting clusters offer actionable insights; for instance, 51% of low-risk firms are found to include the term 'social recruitment' in their textual records. The robustness of ASC is confirmed across multiple clustering algorithms, including k-means, k-medians, and k-medoids, with ΔIntra/Inter < 0.13 and ΔSilhouette Coefficient < 0.02. By bridging spectral clustering theory with heterogeneous data applications, ASC enables the identification of meaningful clusters, such as recruitment-focused SMEs exhibiting a 30% lower default risk, thereby supporting more targeted and effective credit interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。