构建监管问答数据集与系统,提升金融合规信息获取效率
RIRAG: Regulatory Information Retrieval and Answer Generation
- 自动生成问题-段落对,连接监管文本与实际问题
- 基于ADGM法规创建2.7万条问答数据,验证系统有效性
- 专为合规场景设计评估指标,确保答案准确无矛盾
政府监管机构发布的监管文件规定了组织必须遵守的规则、指南和标准。这些文件通常篇幅长、结构复杂且更新频繁,解读难度大,组织需投入大量时间和专业人力以确保持续合规。监管自然语言处理(RegNLP)是一个多学科领域,旨在简化对监管规则与义务的访问与理解。本文提出生成问题-段落对的任务,通过自动生成问题并匹配相关监管段落,推动监管问答系统的发展。我们构建了ObliQA数据集,包含从阿布扎比全球市场(ADGM)金融监管文件中提取的27,869个问题,设计了一个基准的监管信息检索与答案生成(RIRAG)系统,并采用RePASs这一新评估指标,检验生成答案是否准确涵盖所有相关义务且无矛盾。
原文摘要 · Abstract (English)
Regulatory documents, issued by governmental regulatory bodies, establish rules, guidelines, and standards that organizations must adhere to for legal compliance. These documents, characterized by their length, complexity and frequent updates, are challenging to interpret, requiring significant allocation of time and expertise on the part of organizations to ensure ongoing compliance. Regulatory Natural Language Processing (RegNLP) is a multidisciplinary field aimed at simplifying access to and interpretation of regulatory rules and obligations. We introduce a task of generating question-passages pairs, where questions are automatically created and paired with relevant regulatory passages, facilitating the development of regulatory question-answering systems. We create the ObliQA dataset, containing 27,869 questions derived from the collection of Abu Dhabi Global Markets (ADGM) financial regulation documents, design a baseline Regulatory Information Retrieval and Answer Generation (RIRAG) system and evaluate it with RePASs, a novel evaluation metric that tests whether generated answers accurately capture all relevant obligations while avoiding contradictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。