用大模型直接分析信贷文本,自动评估风险
LendNova: Towards Automated Credit Risk Assessment with Language Models
- 用语言模型直接处理原始信贷文本,无需人工特征工程
- 在真实数据上表现优异,提升评估准确性和效率
- 适合金融风控、智能信贷系统研发人员参考
信用风险评估在金融领域至关重要,但传统方法依赖昂贵的特征化模型,难以充分利用原始信贷记录中的全部信息。本文提出LendNova,首个面向信用风险评估的自动化端到端流程,通过先进自然语言处理技术和语言模型,直接利用原始、术语密集的征信文本,学习任务相关表征,无需人工特征工程。该方法自动捕捉文本中嵌入的模式与风险信号,替代传统预处理步骤,降低耗时与成本,提升可扩展性。在真实数据上的评估表明其具备高精度和高效性的潜力。LendNova为智能信用风险代理建立了基准,验证了语言模型在该领域的可行性,并为未来构建更精准、灵活、自动化的金融决策基础系统奠定基础。
原文摘要 · Abstract (English)
Credit risk assessment is essential in the financial sector, but has traditionally depended on costly feature-based models that often fail to utilize all available information in raw credit records. This paper introduces LendNova, the first practical automated end-to-end pipeline for credit risk assessment, designed to utilize all available information in raw credit records by leveraging advanced NLP techniques and language models. LendNova transforms risk modeling by operating directly on raw, jargon-heavy credit bureau text using a language model that learns task-relevant representations without manual feature engineering. By automatically capturing patterns and risk signals embedded in the text, it replaces manual preprocessing steps, reducing costs and improving scalability. Evaluation on real-world data further demonstrates its strong potential in accurate and efficient risk assessment. LendNova establishes a baseline for intelligent credit risk agents, demonstrating the feasibility of language models in this domain. It lays the groundwork for future research toward foundation systems that enable more accurate, adaptable, and automated financial decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。