arXiv:2504.06293q-fin.RMcs.LG2025-04被引 4

用金融监管文档训练模型,提升风险问答系统检索精度。

Generative AI Enhanced Financial Risk Management Information Retrieval

  • 基于OSFI 1991–2024年94份指南构建专用数据集
  • 微调后的RiskEmbed在排名指标上显著优于通用模型
  • 适合金融机构和研究者开发精准的AI风控工具

金融风险管理涉及识别、评估与应对风险以维持稳定并满足合规要求。从大量监管文件中提取相关洞察是一项复杂挑战,需依赖先进的检索与语言模型。本文提出RiskData数据集,专为金融风险领域微调嵌入模型而设计;并构建RiskEmbed模型,旨在提升金融问答系统中的检索准确性。该数据集源自加拿大金融机构监管办公室(OSFI)1991至2024年间发布的94份监管指南。我们对前沿的Sentence BERT嵌入模型进行微调,以增强其在检索增强生成(RAG)系统中的领域适应能力。实验表明,RiskEmbed在排名性能上显著优于通用及金融领域嵌入模型,带来实质性提升。通过开源数据集与模型,本工作为金融机构与研究者提供可复用资源,助力更精准高效的金融风险智能管理解决方案。

原文摘要 · Abstract (English)

Risk management in finance involves recognizing, evaluating, and addressing financial risks to maintain stability and ensure regulatory compliance. Extracting relevant insights from extensive regulatory documents is a complex challenge requiring advanced retrieval and language models. This paper introduces RiskData, a dataset specifically curated for finetuning embedding models in risk management, and RiskEmbed, a finetuned embedding model designed to improve retrieval accuracy in financial question-answering systems. The dataset is derived from 94 regulatory guidelines published by the Office of the Superintendent of Financial Institutions (OSFI) from 1991 to 2024. We finetune a state-of-the-art sentence BERT embedding model to enhance domain-specific retrieval performance typically for Retrieval-Augmented Generation (RAG) systems. Experimental results demonstrate that RiskEmbed significantly outperforms general-purpose and financial embedding models, achieving substantial improvements in ranking metrics. By open-sourcing both the dataset and the model, we provide a valuable resource for financial institutions and researchers aiming to develop more accurate and efficient risk management AI solutions.

金融风险信息检索嵌入模型RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。