用计算方法识别欧盟法律中的监管条文,助力评估法规密度与严格性。
Computational Identification of Regulatory Statements in EU Legislation
- 基于句法依存分析与变压器模型两种方法自动识别监管条文。
- 准确率分别达80%和84%,一致性系数Kappa为0.58。
- 适合关注法律自动化分析、政策评估的研究者使用。
在立法中识别监管条文有助于构建衡量法规密度与严格性的指标。针对1952至2023年间约18万份欧盟法律文件的快速增长,开发计算方法对规模化识别此类条文具有重要意义。现有研究对监管条文的定义宽严不一。本文基于机构语法工具提出明确界定,并比较了两种不同方法:一种依赖依存句法分析,另一种采用基于变压器的机器学习模型。两种方法表现相近,准确率分别为80%和84%,一致系数Kappa为0.58。高准确率但非极高的一致性表明二者优势可互补,具备融合潜力。
原文摘要 · Abstract (English)
Identifying regulatory statements in legislation is useful for developing metrics to measure the regulatory density and strictness of legislation. A computational method is valuable for scaling the identification of such statements from a growing body of EU legislation, constituting approximately 180,000 published legal acts between 1952 and 2023. Past work on extraction of these statements varies in the permissiveness of their definitions for what constitutes a regulatory statement. In this work, we provide a specific definition for our purposes based on the institutional grammar tool. We develop and compare two contrasting approaches for automatically identifying such statements in EU legislation, one based on dependency parsing, and the other on a transformer-based machine learning model. We found both approaches performed similarly well with accuracies of 80% and 84% respectively and a K alpha of 0.58. The high accuracies and not exceedingly high agreement suggests potential for combining strengths of both approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。