用大模型提升银行监管措施的前后一致性。
LLM-based IR-system for Bank Supervisors
- 融合词法、语义与法规模糊匹配,精准检索历史监管案例。
- 在部分标注数据下仍表现稳定,MAP@100达0.83,MRR@100达0.92。
- 适合银行监管人员快速生成合规且一致的监管措施。
银行监管者需确保新出台的监管措施与历史实践保持一致。为此,我们提出一种面向监管场景的信息检索(IR)系统,用于辅助起草一致且有效的监管措施。该系统接收现场检查发现,从全面数据库中检索最相关的历史发现及其对应措施,为新情况提供决策依据。通过结合词法、语义及资本要求条例(CRR)的模糊集匹配技术,确保检索结果与当前案例高度契合。采用蒙特卡洛方法验证系统性能,尤其在部分标注数据条件下展现出强鲁棒性与高准确率。经基于Transformer的去噪自编码器微调后,最终模型在MAP@100上达到0.83,在MRR@100上达到0.92,优于独立的词法模型(如BM25)和语义类BERT模型。
原文摘要 · Abstract (English)
Bank supervisors face the complex task of ensuring that new measures are consistently aligned with historical precedents. To address this challenge, we introduce a novel Information Retrieval (IR) System tailored to assist supervisors in drafting both consistent and effective measures. This system ingests findings from on-site investigations. It then retrieves the most relevant historical findings and their associated measures from a comprehensive database, providing a solid basis for supervisors to write well-informed measures for new findings. Utilizing a blend of lexical, semantic, and Capital Requirements Regulation (CRR) fuzzy set matching techniques, the IR system ensures the retrieval of findings that closely align with current cases. The performance of this system, particularly in scenarios with partially labeled data, is validated through a Monte Carlo methodology, showcasing its robustness and accuracy. Enhanced by a Transformer-based Denoising AutoEncoder for fine-tuning, the final model achieves a Mean Average Precision (MAP@100) of 0.83 and a Mean Reciprocal Rank (MRR@100) of 0.92. These scores surpass those of both standalone lexical models such as BM25 and semantic BERT-like models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。