用大模型把复杂的信贷风险解释变成人能懂的故事。
Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

- 用大模型将XGBoost、GNN等模型的解释转化为通俗叙事
- 专业人员比普通人更严格地检验解释证据的可靠性
- 解释中影响因素名称准确,但方向判断常出错
信贷决策是高风险任务,模型输出需准确且可解释以确保合规。尽管XGBoost和图神经网络(GNNs)提升了预测性能,其解释往往过于技术化,导致利益相关方理解断层,影响审批、拒贷及公平性判断。本文研究大语言模型(LLMs)是否可作为解释层,将事后解释结果转化为适合利益相关方的风险叙述。基于弗雷迪·麦克单户贷款级数据,构建三种管道:标准表格型(XGBoost + SHAP)、纯网络型(GNN + GNNExplainer)和双模态型(结合表格与网络数据)。采用三种LLM配置生成叙述:小规模微调模型(Gemma 3 4B)、大规模微调模型(DeepSeek R1 70B)和零样本商用模型(Gemini 2.5)。通过自动化评估与人类实验(对比信用专业人士与非专业人士在八个决策相关维度的表现),发现:第一,不同管道导致的证据可信度差异大于模型选择的影响,说明解释质量受限于原始证据表示;第二,叙述能可靠识别关键影响因素,但在判断影响方向上可靠性较低,可能影响不利行动通知;第三,专业人士对证据的标准更严。研究对风险模型治理提出建议,包括部署考量及领域对齐大模型在监管信贷环境中的价值。
原文摘要 · Abstract (English)
Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。