用超图神经网络提升法律文书实体关系抽取效果
Knowledge Augmented Entity and Relation Extraction for Legal Documents with Hypergraph Neural Network
- 基于邻域打包与双仿射机制生成候选实体片段
- 融合司法知识库与多罪并罚等场景构建超图结构
- 在CAIL2022数据集上超越现有模型,适合法律AI研究者
随着中国司法机构数字化进程推进,大量电子法律文档信息积累。为挖掘其价值,法律文书中的实体与关系抽取成为关键任务。然而现有方法普遍缺乏领域知识,难以捕捉司法领域的特殊性。本文提出基于超图神经网络的法律知识增强实体关系抽取方法(Legal-KAHRE),针对毒品相关判决书设计。首先,采用基于邻域导向打包策略与双仿射机制的候选片段生成器,识别可能包含实体的文本段;其次,构建司法领域知识词典,并通过多头注意力机制融入文本编码表示;此外,将共同犯罪、数罪并罚等典型司法案例纳入超图结构设计;最后,利用超图神经网络通过消息传递实现高阶推理。在CAIL2022信息抽取数据集上的实验表明,该方法显著优于现有基线模型。
原文摘要 · Abstract (English)
With the continuous progress of digitization in Chinese judicial institutions, a substantial amount of electronic legal document information has been accumulated. To unlock its potential value, entity and relation extraction for legal documents has emerged as a crucial task. However, existing methods often lack domain-specific knowledge and fail to account for the unique characteristics of the judicial domain. In this paper, we propose an entity and relation extraction algorithm based on hypergraph neural network (Legal-KAHRE) for drug-related judgment documents. Firstly, we design a candidate span generator based on neighbor-oriented packing strategy and biaffine mechanism, which identifies spans likely to contain entities. Secondly, we construct a legal dictionary with judicial domain knowledge and integrate it into text encoding representation using multi-head attention. Additionally, we incorporate domain-specific cases like joint crimes and combined punishment for multiple crimes into the hypergraph structure design. Finally, we employ a hypergraph neural network for higher-order inference via message passing. Experimental results on the CAIL2022 information extraction dataset demonstrate that our method significantly outperforms existing baseline models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。