清理并结构化25万份德语判决书,划分关键法律部分。
Segmentation and Processing of German Court Decisions from Open Legal Data
- 基于随机抽样验证,系统分割判决书中的事实、理由和裁决三部分。
- 构建包含251,038份判决书的结构化数据集,准确率达95%以上。
- 适合法律AI研究者用于检索、引用分析等下游任务。
结构化法律数据对推动德语法律系统的自然语言处理技术至关重要。当前广泛使用的开源数据集Open Legal Data提供了大规模德语判决书集合,但其文本格式不一致,关键部分缺乏明确标记。可靠地分离这些部分不仅有助于语篇角色分类,也支持后续的检索与引用分析。本文从官方Open Legal Data数据集中提取并清洗出251,038份德语判决书,系统性地区分了三个核心部分:裁决(Tenor)、案件事实(Tatbestand)和裁判理由(Entscheidungsgründe),这些在原始数据中常表述不一。为确保提取可靠性,采用置信度95%、误差5%的Cochran公式抽取384个样本,经人工验证确认三部分识别准确。同时将上诉告知(Rechtsmittelbelehrung)作为独立字段提取,因其属程序性说明而非判决内容。最终数据集以JSONL格式公开,可直接用于德语法律系统相关研究。
原文摘要 · Abstract (English)
The availability of structured legal data is important for advancing Natural Language Processing (NLP) techniques for the German legal system. One of the most widely used datasets, Open Legal Data, provides a large-scale collection of German court decisions. While the metadata in this raw dataset is consistently structured, the decision texts themselves are inconsistently formatted and often lack clearly marked sections. Reliable separation of these sections is important not only for rhetorical role classification but also for downstream tasks such as retrieval and citation analysis. In this work, we introduce a cleaned and sectioned dataset of 251,038 German court decisions derived from the official Open Legal Data dataset. We systematically separated three important sections in German court decisions, namely Tenor (operative part of the decision), Tatbestand (facts of the case), and Entscheidungsgründe (judicial reasoning), which are often inconsistently represented in the original dataset. To ensure the reliability of our extraction process, we used Cochran's formula with a 95% confidence level and a 5% margin of error to draw a statistically representative random sample of 384 cases, and manually verified that all three sections were correctly identified. We also extracted the Rechtsmittelbelehrung (appeal notice) as a separate field, since it is a procedural instruction and not part of the decision itself. The resulting corpus is publicly available in the JSONL format, making it an accessible resource for further research on the German legal system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。