构建可动态建模法律概念的标注系统,提升法律文本结构化处理能力。
Annotating Topical Legal Insights from Case Proceedings

- 提出LeDA系统,支持网页端动态创建新标签进行法律概念标注
- 3名评估员在印度最高法院案例中成功标注法律概念,构建概念袋表示
- 适用于无预设本体的法律概念发现场景,适合法律信息检索与判例预测
本文聚焦于从法律案件审理记录中提取概念或主题,因为采用结构化表示而非简单的词袋模型,能显著提升法律文档处理能力。为此,我们提出了一套针对法律案件审理记录的概念体系,并开发了名为LeDA的法律数据标注系统。该系统通过网页界面提供通用的实体或概念标注与裁定功能。其创新之处在于支持动态创建新标签,特别适用于缺乏预定义本体的场景——即标注者在持续审阅文档过程中逐步发现需标注的概念。当前系统已用于对一系列法律文档进行概念标注,构建以概念为单位的语义表示,可用于后续任务如判例检索、判决预测等。此外,本文还描述了3名评估员如何使用LeDA对印度最高法院案件审理记录中的法律概念名称进行标注与裁定。
原文摘要 · Abstract (English)
In this paper, we mainly concentrate on finding concepts or topics from the legal case proceedings, since adopting a structured representation for legal documents, as opposed to a mere bag-of-words flat text representation, can significantly enhance processing capabilities. To achieve this objective, we put forward a set of diverse concepts for legal case proceedings. With this motivation, we propose LeDA, a system for Legal Data Annotation. The system offers the generic functionality of annotating and adjudicating entities or concepts within documents via a web-based interface. A novel feature of our system is that it allows to dynamic create new tags for annotation, which is a particularly useful provision for situations where there exists no pre-defined ontology for the entities (concepts) that need to be annotated - these being rather discovered by annotators as they continue examining more documents. The system that we demonstrate is currently in use to annotate a set of concepts from legal documents to construct semantic representations of documents as bags of concepts that can then be used for several downstream tasks, such as prior case retrieval, judgment prediction, and so on. Along with the system features in general, we also describe how LeDA was used by 3 assessors to annotate and adjudicate legal concept names from Indian Supreme Court case proceedings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。