arXiv:2410.09527cs.CL2024-10中稿 · EMNLP被引 19

首个面向英文法律摘要的基准与模型,解决法律文本生成评估难题。

LexSumm and LexT5: Benchmarking and Modeling Legal Summarization Tasks in English

  • 构建涵盖美英欧印的8个法律摘要数据集,填补领域空白。
  • 零样本测试显示大模型摘要仍存抽象与忠实性错误。
  • 发布专为法律设计的序列到序列模型LexT5,提升生成质量。

在不断发展的自然语言处理领域,基准测试是衡量进展的重要标尺。然而,现有法律NLP基准仅关注预测任务,忽视生成任务。本文构建了LexSumm,一个面向英文法律摘要任务的基准,包含来自美国、英国、欧盟和印度等不同司法管辖区的8个英文法律摘要数据集。同时,我们发布了LexT5,一种面向法律领域的序列到序列模型,以克服现有BERT风格编码器模型在法律领域生成能力不足的问题。通过在LegalLAMA上进行零样本探测和在LexSumm上微调,评估其性能。分析发现,即使是零样本大模型生成的摘要也存在抽象和忠实性错误,表明仍有改进空间。LexSumm基准与LexT5模型已开源,地址为https://github.com/TUMLegalTech/LexSumm-LexT5。

原文摘要 · Abstract (English)

In the evolving NLP landscape, benchmarks serve as yardsticks for gauging progress. However, existing Legal NLP benchmarks only focus on predictive tasks, overlooking generative tasks. This work curates LexSumm, a benchmark designed for evaluating legal summarization tasks in English. It comprises eight English legal summarization datasets, from diverse jurisdictions, such as the US, UK, EU and India. Additionally, we release LexT5, legal oriented sequence-to-sequence model, addressing the limitation of the existing BERT-style encoder-only models in the legal domain. We assess its capabilities through zero-shot probing on LegalLAMA and fine-tuning on LexSumm. Our analysis reveals abstraction and faithfulness errors even in summaries generated by zero-shot LLMs, indicating opportunities for further improvements. LexSumm benchmark and LexT5 model are available at https://github.com/TUMLegalTech/LexSumm-LexT5.

法律AI摘要生成基准测试序列生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。