arXiv:2512.09434cs.CLcs.AI2025-12中稿 · JURIX AI4A2J Works…

构建德语法院判决摘要数据集,助力大模型生成通俗易懂的司法新闻。

CourtPressGER: A German Court Decision to Press Release Summarization Dataset

  • 构建6.4k条三元组数据:判决书、人工撰写的新闻稿、LLM生成提示
  • 大模型可生成高质量摘要,小模型需分层结构处理长文本
  • 适合司法AI、自然语言生成与可解释性研究者使用

德国最高法院发布的官方新闻稿面向公众及专业群体阐释司法裁决。以往NLP研究侧重技术性摘要,忽视公众传播需求。本文提出CourtPressGER,一个包含6.4k个三元组的数据集,包含判决文书、人工撰写的新闻稿以及用于引导大模型生成类似稿件的合成提示。该基准可用于训练和评估大模型从长篇判决中生成准确且易读摘要的能力。通过参考指标、事实一致性检测、大模型评分及专家排序进行多维度评测。结果显示,大模型在生成质量上表现优异,层级结构损失小;小模型则需采用分层架构处理长判决。初步测试表明,人工撰写的新闻稿质量最高。

原文摘要 · Abstract (English)

Official court press releases from Germany's highest courts present and explain judicial rulings to the public, as well as to expert audiences. Prior NLP efforts emphasize technical headnotes, ignoring citizen-oriented communication needs. We introduce CourtPressGER, a 6.4k dataset of triples: rulings, human-drafted press releases, and synthetic prompts for LLMs to generate comparable releases. This benchmark trains and evaluates LLMs in generating accurate, readable summaries from long judicial texts. We benchmark small and large LLMs using reference-based metrics, factual-consistency checks, LLM-as-judge, and expert ranking. Large LLMs produce high-quality drafts with minimal hierarchical performance loss; smaller models require hierarchical setups for long judgments. Initial benchmarks show varying model performance, with human-drafted releases ranking highest.

司法AI摘要生成大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。