构建德语法律条文逻辑结构标注数据集,助力自动法律推理
ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts

- 标注德语法律条文中的构成要件与法律后果,实现细粒度结构识别
- 涵盖三类法律典籍,验证模型在跨法典场景下的泛化能力
- BERT与大模型表现最优,为法律文本分析提供新基准
法律文本的自动结构分析是法律科技的核心挑战之一,而其逻辑成分的提取仍面临巨大困难。本文提出识别并分割德语法律条文中的构成要件(Tatbestand)与法律后果(Rechtsfolge)的任务。为此,我们构建了 ANNOTARES(Tatbestand-Rechtsfolge 序列标注数据集),包含三类不同法律典籍的德语法律文本,采用跨度级标注。该数据集旨在评估模型在特定领域内的性能及跨法典的泛化能力。我们对多种架构进行了基准测试:基于规则的基线、CRFs、BiLSTMs、BiLSTM-CRF,以及现代 Transformer 模型(包括 BERT 变体和大语言模型方法)。结果表明,BERT 和大语言模型在捕捉法律语言复杂句法结构方面表现更优。我们已公开该数据集,以推动自动化法律推理研究。
原文摘要 · Abstract (English)
The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a significant challenge. In this paper, we introduce the task of identifying and segmenting legal conditions (Tatbestand) and legal consequences (Rechtsfolge) within German statutory texts. To support this task, we present ANNOTARES (Annotations of Tatbestand-Rechtsfolge Sequences), a novel dataset comprising German law texts with span-level annotations. Spanning three distinct legal codes, the dataset is designed to evaluate both domain-specific performance and cross-statute generalizability. We benchmark diverse architectural approaches: a rule-based baseline, CRFs, BiLSTMs, BiLSTM-CRF, and modern Transformer-based models, including BERT variants and LLM-based methods. Our results demonstrate that BERT and LLM-based models achieve superior performance in capturing the complex syntactic structures of legal language. We release our dataset to facilitate further research in automated legal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。