arXiv:2502.05836cs.CLcs.AI2025-02NAACL被引 16

构建首个印度判例文结构化数据集,提升法律文本语义解析能力

LegalSeg: Unlocking the Structure of Indian Legal Judgments Through Rhetorical Role Classification

  • 基于7种修辞角色对超7000份判决书进行标注,建立大规模语料库
  • 结合上下文与句间关系的模型显著优于仅依赖单句特征的方法
  • 适合法律AI研究者、司法数字化推进者参考

本文聚焦印度判例文的语义分割任务,提出面向修辞角色分类的新方法。构建了LegalSeg数据集,包含超过7,000份文档、140万条句子,每条标注7种修辞角色。在该数据集上评估了多种先进模型,包括分层BiLSTM-CRF、ToInLegalBERT、图神经网络(GNN)及角色感知Transformer,还引入探索性模型RhetoricLLaMA。结果表明,融合更广上下文、结构关系和序列信息的模型表现更优。通过利用邻近句子的预测或真实标签,进一步验证了上下文作用。尽管如此,相似角色区分与类别不平衡问题仍存。本研究展示了先进方法在法律文本理解中的潜力,并为法律NLP发展奠定基础。

原文摘要 · Abstract (English)

In this paper, we address the task of semantic segmentation of legal documents through rhetorical role classification, with a focus on Indian legal judgments. We introduce LegalSeg, the largest annotated dataset for this task, comprising over 7,000 documents and 1.4 million sentences, labeled with 7 rhetorical roles. To benchmark performance, we evaluate multiple state-of-the-art models, including Hierarchical BiLSTM-CRF, TransformerOverInLegalBERT (ToInLegalBERT), Graph Neural Networks (GNNs), and Role-Aware Transformers, alongside an exploratory RhetoricLLaMA, an instruction-tuned large language model. Our results demonstrate that models incorporating broader context, structural relationships, and sequential sentence information outperform those relying solely on sentence-level features. Additionally, we conducted experiments using surrounding context and predicted or actual labels of neighboring sentences to assess their impact on classification accuracy. Despite these advancements, challenges persist in distinguishing between closely related roles and addressing class imbalance. Our work underscores the potential of advanced techniques for improving legal document understanding and sets a strong foundation for future research in legal NLP.

法律AI语义分割判例分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。