用大模型构建意大利立法流程数据集,助力法律流程挖掘研究
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
- 用大语言模型从非结构化文本中提取立法事件,构建全流程数据
- 覆盖1987至2022年意大利议会立法过程,含数万条事件记录
- 为法律流程挖掘提供首个高质量基准数据集,适合法律与AI交叉研究者
流程挖掘(PM)最初用于工业和商业场景,近年来被引入社会系统,包括法律领域。然而,其在法律领域的应用受限于数据的可获取性和质量。本文提出ProLiFIC(意大利议院程序性立法流程),一个涵盖1987至2022年意大利立法过程的综合性事件日志。该数据集基于来自Normattiva门户的非结构化文本,通过大型语言模型(LLMs)进行结构化处理,契合当前将流程挖掘与大模型结合的趋势。我们展示了初步分析案例,并建议将ProLiFIC作为法律流程挖掘的基准,推动该领域的新发展。
原文摘要 · Abstract (English)
Process Mining (PM), initially developed for industrial and business contexts, has recently been applied to social systems, including legal ones. However, PM's efficacy in the legal domain is limited by the accessibility and quality of datasets. We introduce ProLiFIC (Procedural Lawmaking Flow in Italian Chambers), a comprehensive event log of the Italian lawmaking process from 1987 to 2022. Created from unstructured data from the Normattiva portal and structured using large language models (LLMs), ProLiFIC aligns with recent efforts in integrating PM with LLMs. We exemplify preliminary analyses and propose ProLiFIC as a benchmark for legal PM, fostering new developments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。