构建专家标注的刑法四要件知识库,提升法律大模型推理准确性
JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoning
- 基于法律权威源和多元解释方法构建层级化四要件标注体系
- 覆盖155个罪名,显著提升法律案例检索与相似罪名辨析性能
- 适合法律AI研究者、司法科技开发者使用
近年来,大语言模型(LLMs)被广泛应用于法律任务。为增强其对法律文本的理解并提高推理准确性,融入法律理论成为重要方向。其中,四要件理论(FET)通过主体、客体、主观方面和客观方面定义犯罪构成,应用最广。尽管已有研究尝试引导LLM遵循FET,但我们的评估显示,模型生成的四要件常不完整且代表性不足,限制了其在法律推理中的效果。为此,我们提出JUREX-4E,一个涵盖155个刑事罪名的专家标注四要件知识库。标注采用基于法律来源有效性的渐进式层级框架,并融合多种解释方法以确保精确性与权威性。我们在相似罪名辨析任务上评估JUREX-4E,并应用于法律案例检索。实验结果验证了其高质量及对下游任务的显著提升,凸显其在推动法律AI发展中的潜力。数据集与代码已公开于:https://github.com/THUlawtech/JUREX
原文摘要 · Abstract (English)
In recent years, Large Language Models (LLMs) have been widely applied to legal tasks. To enhance their understanding of legal texts and improve reasoning accuracy, a promising approach is to incorporate legal theories. One of the most widely adopted theories is the Four-Element Theory (FET), which defines the crime constitution through four elements: Subject, Object, Subjective Aspect, and Objective Aspect. While recent work has explored prompting LLMs to follow FET, our evaluation demonstrates that LLM-generated four-elements are often incomplete and less representative, limiting their effectiveness in legal reasoning. To address these issues, we present JUREX-4E, an expert-annotated four-element knowledge base covering 155 criminal charges. The annotations follow a progressive hierarchical framework grounded in legal source validity and incorporate diverse interpretive methods to ensure precision and authority. We evaluate JUREX-4E on the Similar Charge Disambiguation task and apply it to Legal Case Retrieval. Experimental results validate the high quality of JUREX-4E and its substantial impact on downstream legal tasks, underscoring its potential for advancing legal AI applications. The dataset and code are available at: https://github.com/THUlawtech/JUREX
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。