arXiv:2503.12100cs.CL2025-03

构建波兰议会立法文本数据集,评估大模型在法律文本分析中的表现。

Large Language Models in Legislative Content Analysis: A Dataset from the Polish Parliament

  • 基于波兰官方立法网站构建法律文本数据集
  • 验证大模型在法律语境理解上的潜力与挑战
  • 适合关注法律AI与自然语言处理的研究者

大型语言模型(LLMs)是处理自然语言的优秀方法之一,部分原因在于其通用性。同时,领域专用的LLM在实际应用中更具可行性。本文通过官方立法机构网站获取数据,构建了一个新的自然语言数据集。研究聚焦于提出三项自然语言处理任务,以评估LLMs在波兰法律体系下的立法内容分析效果。关键发现表明,LLMs在自动化和提升立法内容分析方面具有潜力,但对法律语境的理解仍存在特定挑战。该研究推动了自然语言处理在法律领域的进展,特别是在波兰语环境下的应用。研究表明,即使是常见的公开数据,也可有效用于立法内容分析。

原文摘要 · Abstract (English)

Large language models (LLMs) are among the best methods for processing natural language, partly due to their versatility. At the same time, domain-specific LLMs are more practical in real-life applications. This work introduces a novel natural language dataset created by acquired data from official legislative authorities' websites. The study focuses on formulating three natural language processing (NLP) tasks to evaluate the effectiveness of LLMs on legislative content analysis within the context of the Polish legal system. Key findings highlight the potential of LLMs in automating and enhancing legislative content analysis while emphasizing specific challenges, such as understanding legal context. The research contributes to the advancement of NLP in the legal field, particularly in the Polish language. It has been demonstrated that even commonly accessible data can be practically utilized for legislative content analysis.

法律AI大模型波兰语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。