用大模型生成信贷风险解释,能还原重要特征排序但自主解释效果有限。
Could Large Language Models work as Post-hoc Explainability Tools in Credit Risk Models?
- 用可控提示让大模型复现特征重要性排序
- 自主生成解释时与基准方法对齐度低
- 适合做可读性接口,不适合作为决策依据
大语言模型(LLMs)在将模型解释转化为人类可读叙述方面展现出潜力。本研究评估了LLMs作为信贷风险模型后验可解释性工具的可行性,重点关注其保持特征重要性排序及自主生成解释的能力。基于LendingClub数据集,我们对比了GPT-4-turbo、Claude-Sonnet-4.5和Gemini-2.5-Flash三款主流大模型在三种情形下的表现:控制提示下与SHAP及系数法归因结果的匹配度。结果显示,在受控提示条件下,LLMs能可靠再现参考排序;但在自主生成解释时,与基准方法的对齐程度显著下降。研究建议,应将LLMs定位为可解释性叙述接口,而非正式归因方法的替代品,尤其在信贷风险管理中需谨慎使用。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise in translating model-based explanations into human-readable narratives. This study evaluates whether LLMs can serve as post-hoc explainability interfaces for credit risk models, focusing on their ability to preserve feature-importance rankings and generate autonomous explanations. Using a LendingClub dataset, we compare LLM outputs with SHAP and coefficient-based attributions on three major LLMs, including GPT-4-turbo, Claude-Sonnet-4.5, and Gemini-2.5-Flash. Results indicate that LLMs reliably reproduce reference rankings under controlled prompts but show limited alignment when generating explanations autonomously. These findings suggest that LLMs are best deployed as narrative interfaces rather than substitutes for formal attribution methods in credit risk governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。