arXiv:2605.24534cs.CL2026-05

用法院判例自动生成法律注释,无需人工框架

Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation

论文配图:Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation
图 1 · 摘自论文原文
  • 从判例中提取段落、聚类并生成标题与内容
  • 生成的注释在5个维度上表现良好,可分钟级更新
  • 适合法律AI研究者和司法自动化从业者

我们提出一个全自动流程,将大量法院判决转化为针对法条的法律注释,无需人工构建教义框架。基于4555份引用德国《民法典》(BGB)第242、280、812及823条的德国联邦最高法院判决,我们提取段落级文本,总结推理逻辑并提取关键词,经嵌入与聚类处理。每个聚类由大语言模型生成标题并合成带引用的内容,再由四个先进LLM整合为连贯注释。评估涵盖五个维度:主题相关性、标题匹配度、引文忠实度、聚类区分度与逻辑顺序,采用人类专家与LLM裁判双重验证。结果表明,从判例中自动挖掘类似注释并生成可快速更新的报告是可行的,成本极低,但受限于数据来源和法律推理的规范性要求。

原文摘要 · Abstract (English)

We present a fully automated pipeline that transforms large collections of court decisions into legal commentaries for statutes - without providing any handcrafted doctrinal framework. Using 4.555 decisions of the German Federal Court of Justice that cite sections 242, 280, 812 and 823 of the German Civil Code (BGB), we extract paragraph-level chunks, summarize their reasoning, and derive keywords, which are embedded and clustered. For each cluster, an LLM generates headings and synthesizes citation-rich sections, which are then merged into coherent commentaries by four state-of-the-art LLMs. We evaluate along five dimensions - topical relevance, heading-match, citation faithfulness, cluster distinction and logical ordering - using both a human expert and an LLM-judge. Our results show that commentary-like argument mining from court decisions to generate reports that can be refreshed within minutes at minimal cost is feasible, yet they highlight limitations arising from restricted sources and the normativity of legal reasoning.

法律AI生成模型判例分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。