arXiv:2501.10677cs.LGcs.AI2025-01

让大模型更适配金融数据,提升小样本信用评分准确率

Class-Imbalanced-Aware Adaptive Dataset Distillation for Scalable Pretrained Model on Credit Scoring

  • 用数据蒸馏技术让预训练模型适应小规模金融表格数据
  • 考虑类别不平衡,使模型在真实金融数据上AUC提升2.5%
  • 适合想用大模型做信用评分的金融从业者和研究者

人工智能显著提升了信用评分技术。尽管深度学习模型表现优异,主流仍倾向使用树结构模型,因其在表格数据上预测能力强。尽管预训练模型发展迅速,其在金融领域的应用多限于问答任务,用于表格型信用评分数据的研究仍较少。表格式大模型如TabPFN虽使大模型应用于信用评分成为可能,但仅能处理有限样本。本文提出新框架,结合表格式数据蒸馏技术与预训练模型,提升TabPFN的可扩展性。此外,金融数据普遍存在类别不平衡,但其在数据蒸馏中的影响未被研究。本文在蒸馏过程中引入不平衡感知机制,显著提升金融数据上的性能(如AUC提升2.5%)。本研究为大模型在金融表格数据上的应用提供了新路径,并对比分析了类别不平衡对蒸馏过程的影响,有望拓展大模型在金融领域的下游任务应用。

原文摘要 · Abstract (English)

The advent of artificial intelligence has significantly enhanced credit scoring technologies. Despite the remarkable efficacy of advanced deep learning models, mainstream adoption continues to favor tree-structured models due to their robust predictive performance on tabular data. Although pretrained models have seen considerable development, their application within the financial realm predominantly revolves around question-answering tasks and the use of such models for tabular-structured credit scoring datasets remains largely unexplored. Tabular-oriented large models, such as TabPFN, has made the application of large models in credit scoring feasible, albeit can only processing with limited sample sizes. This paper provides a novel framework to combine tabular-tailored dataset distillation technique with the pretrained model, empowers the scalability for TabPFN. Furthermore, though class imbalance distribution is the common nature in financial datasets, its influence during dataset distillation has not been explored. We thus integrate the imbalance-aware techniques during dataset distillation, resulting in improved performance in financial datasets (e.g., a 2.5% enhancement in AUC). This study presents a novel framework for scaling up the application of large pretrained models on financial tabular datasets and offers a comparative analysis of the influence of class imbalance on the dataset distillation process. We believe this approach can broaden the applications and downstream tasks of large models in the financial domain.

信用评分大模型数据蒸馏类别不平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。