融合财务数据与文本,提升债券违约预测的准确与可解释性
Why Bonds Fail Differently? Explainable Multimodal Learning for Multi-Class Default Prediction
- 结合时间序列与文本数据,用时序感知LSTM处理不规则金融数据
- 在1994家企业上实现召回率、F1值等指标超越传统与深度学习模型
- 注意力分析揭示经济直观的违约原因,适合金融风控与监管决策
近年来,中国债券市场在监管改革与宏观经济波动背景下违约事件激增。传统机器学习难以捕捉金融数据的不规则性与时间依赖性,而多数深度学习模型缺乏可解释性,不利于金融决策。为此,我们提出EMDLOT(可解释多模态时序深度学习),用于多类别债券违约预测。该框架融合数值型时间序列(财务/宏观经济指标)与非结构化文本数据(债券募集说明书),采用时序感知LSTM处理不规则序列,并通过软聚类与多层级注意力机制增强可解释性。在1994家中国企业(2015–2024年)上的实验表明,EMDLOT在召回率、F1-score与mAP等指标上优于传统(如XGBoost)与深度学习(如LSTM)基准,尤其在识别违约及延期企业方面表现突出。消融实验证明各组件有效性,注意力分析揭示了经济上合理的违约驱动因素。本工作提供了一套实用工具与可信的透明金融风险建模框架。
原文摘要 · Abstract (English)
In recent years, China's bond market has seen a surge in defaults amid regulatory reforms and macroeconomic volatility. Traditional machine learning models struggle to capture financial data's irregularity and temporal dependencies, while most deep learning models lack interpretability-critical for financial decision-making. To tackle these issues, we propose EMDLOT (Explainable Multimodal Deep Learning for Time-series), a novel framework for multi-class bond default prediction. EMDLOT integrates numerical time-series (financial/macroeconomic indicators) and unstructured textual data (bond prospectuses), uses Time-Aware LSTM to handle irregular sequences, and adopts soft clustering and multi-level attention to boost interpretability. Experiments on 1994 Chinese firms (2015-2024) show EMDLOT outperforms traditional (e.g., XGBoost) and deep learning (e.g., LSTM) benchmarks in recall, F1-score, and mAP, especially in identifying default/extended firms. Ablation studies validate each component's value, and attention analyses reveal economically intuitive default drivers. This work provides a practical tool and a trustworthy framework for transparent financial risk modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。