arXiv:2506.06619cs.CL2025-06ACL被引 7

构建法律文书辅助基准,评估大模型写作与推理能力

BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs

  • 设计三类法律文书任务:论点摘要、论点补全、判例检索
  • 大模型在摘要和引导补全上超越人类,但判例检索表现差
  • 为法律AI研究提供真实场景评测工具,适合法律科技开发者

法律工作中尚未被充分探索的核心环节是法律文书的撰写与修改。这不仅需要对管辖地法律(包括判例与法规)有深入理解,还需具备提出新论点、拓展法律边界并创作说服性论证的能力。为评估语言模型在这些法律技能上的表现,我们引入BRIEFME,一个专注于法律文书的新数据集。该数据集包含三项任务:论点摘要、论点补全与判例检索,旨在帮助法律专业人士撰写文书。本文详细描述了任务构建过程,进行了分析,并展示了当前模型的表现。结果显示,现有大型语言模型(LLMs)在摘要和引导式补全任务中已表现优异,甚至优于人工生成的标题;但在真实论点补全和相关判例检索任务上表现不佳。我们期望该数据集能推动法律自然语言处理领域的进一步发展,以切实支持法律实务工作。

原文摘要 · Abstract (English)

A core part of legal work that has been under-explored in Legal NLP is the writing and editing of legal briefs. This requires not only a thorough understanding of the law of a jurisdiction, from judgments to statutes, but also the ability to make new arguments to try to expand the law in a new direction and make novel and creative arguments that are persuasive to judges. To capture and evaluate these legal skills in language models, we introduce BRIEFME, a new dataset focused on legal briefs. It contains three tasks for language models to assist legal professionals in writing briefs: argument summarization, argument completion, and case retrieval. In this work, we describe the creation of these tasks, analyze them, and show how current models perform. We see that today's large language models (LLMs) are already quite good at the summarization and guided completion tasks, even beating human-generated headings. Yet, they perform poorly on other tasks in our benchmark: realistic argument completion and retrieving relevant legal cases. We hope this dataset encourages more development in Legal NLP in ways that will specifically aid people in performing legal work.

法律AI文本生成判例检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。