构建跨司法管辖区AI法规问答系统,提升政策与法律工作者检索效率
Navigating Global AI Regulation: A Multi-Jurisdictional Retrieval-Augmented Generation System

- 按法律类型分块处理异构法规文本,保留结构信息
- 通过实体识别与元数据路由,精准定位法律引用来源
- 优先重排已立法文件,确保权威内容优先输出
跨司法管辖区的AI监管日益复杂,给政策制定者、法律从业者和研究人员带来挑战。为此,我们提出一种多司法管辖区检索增强生成系统,用于全球AI监管信息的高效获取。系统涵盖68个司法管辖区的242份文档,包括欧盟人工智能法案等正式立法,以及各国人工智能战略等非结构化政策文件。技术上实现三项创新:针对法律类型的分块策略,保持异构文档中的法律结构;基于实体检测与元数据的条件检索路由机制,支持法律引文精准定位;优先级重排序策略,提升已立法文本在结果中的权重。对50个查询的评估显示,系统在单实体和多司法管辖区比较类问题上均表现优异,平均忠实度达0.87,平均相关性达0.84。其中单实体查询的忠实度为0.86,相关性为0.92;多司法管辖区比较查询的忠实度为0.88,相关性为0.75。结果表明,领域专用的检索策略在应对复杂异构监管语料方面具有显著有效性。
原文摘要 · Abstract (English)
Navigating AI regulation across jurisdictions is increasingly difficult for policymakers, legal professionals, and researchers. To address this, we present a multi-jurisdictional Retrieval-Augmented Generation system for global AI regulation. Our corpus includes 242 documents across 68 jurisdictions, ranging from formal legislation like the EU AI Act to unstructured policy documents such as national AI strategies. The system makes three technical contributions: type-specific chunking that preserve legal structure across heterogenous documents; conditional retrieval routing with entity detection and metadata for legal citations; and priority-based re-ranking to boost enacted legislation over policy and secondary sources. Evaluation of 50 queries reveals strong performance across both single-entity and multi-jurisdictional questions, achieving 0.87 average faithfulness and 0.84 average answer relevancy. Single-entity queries achieve 0.86 average faithfulness and 0.92 average answer relevancy, while multi-jurisdictional comparison queries achieve 0.88 average faithfulness and 0.75 average answer relevancy. These findings highlight the effectiveness of domain-specific retrieval strategies for navigating complex, heterogenous regulatory corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。