融合年报文本与财务数据,提升企业信用评级准确率8-12%
CreditARF: A Framework for Corporate Credit Rating with Annual Report and Financial Feature Integration
- 用FinBERT提取年报文本特征,与财务数据联合建模
- 在自建大尺度数据集上实现8-12%的准确率提升
- 适合关注金融文本挖掘与信用评估的研究者
企业信用评级是市场经济中的关键中介服务,在维护经济秩序中起重要作用。现有评级模型依赖财务指标与深度学习,但常忽视非财务数据(如企业年报)中的信息。本文提出一种融合财务数据与年报文本特征的信用评级框架,利用FinBERT提取年报中的语义特征,实现多源信息整合。同时构建了大规模数据集CCRD,包含传统财务数据与年报文本数据。实验表明,该方法使评级预测准确率提升8%-12%,显著增强信用评级的有效性与可靠性。
原文摘要 · Abstract (English)
Corporate credit rating serves as a crucial intermediary service in the market economy, playing a key role in maintaining economic order. Existing credit rating models rely on financial metrics and deep learning. However, they often overlook insights from non-financial data, such as corporate annual reports. To address this, this paper introduces a corporate credit rating framework that integrates financial data with features extracted from annual reports using FinBERT, aiming to fully leverage the potential value of unstructured text data. In addition, we have developed a large-scale dataset, the Comprehensive Corporate Rating Dataset (CCRD), which combines both traditional financial data and textual data from annual reports. The experimental results show that the proposed method improves the accuracy of the rating predictions by 8-12%, significantly improving the effectiveness and reliability of corporate credit ratings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。