构建首个覆盖1815年以来美国最高法院判例的长文本摘要数据集
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
- 收集2.56万份最高法院判决与官方摘要,构建最大公开法律摘要数据集
- 小模型Mistral 7b在自动指标上胜过大模型,但人类评估发现其存在幻觉
- 强调人工评估在高风险领域的重要性,揭示自动评估的局限性
本文提出CaseSumm,一个面向法律领域的长文本摘要数据集,旨在填补现有摘要评估数据集在长度与复杂性上的不足。我们收集了2.56万份美国最高法院(SCOTUS)判决及其官方摘要(即syllabuses),该数据集是目前最大的公开法律案例摘要数据集,且首次涵盖自1815年以来的判决摘要。我们对大语言模型生成的摘要进行了全面评估,结合自动指标与专家人工评价,结果显示:尽管小规模开源模型Mistral 7b在多数自动指标上优于大型模型,且能生成类似官方摘要的文本,但专家指出其存在明显幻觉;相比之下,GPT-4生成的摘要在清晰度、敏感性与准确性方面更受认可。进一步分析表明,基于LLM的评估与人工评估的相关性并不优于传统自动指标。此外,我们识别出生成摘要中的具体错误,包括先例引用错误与案件事实误述。这些发现揭示了当前自动评估方法在法律摘要任务中的局限性,凸显了人工评估在复杂高风险领域中的关键作用。CaseSumm已开放获取于https://huggingface.co/datasets/ChicagoHAI/CaseSumm。
原文摘要 · Abstract (English)
This paper introduces CaseSumm, a novel dataset for long-context summarization in the legal domain that addresses the need for longer and more complex datasets for summarization evaluation. We collect 25.6K U.S. Supreme Court (SCOTUS) opinions and their official summaries, known as "syllabuses." Our dataset is the largest open legal case summarization dataset, and is the first to include summaries of SCOTUS decisions dating back to 1815. We also present a comprehensive evaluation of LLM-generated summaries using both automatic metrics and expert human evaluation, revealing discrepancies between these assessment methods. Our evaluation shows Mistral 7b, a smaller open-source model, outperforms larger models on most automatic metrics and successfully generates syllabus-like summaries. In contrast, human expert annotators indicate that Mistral summaries contain hallucinations. The annotators consistently rank GPT-4 summaries as clearer and exhibiting greater sensitivity and specificity. Further, we find that LLM-based evaluations are not more correlated with human evaluations than traditional automatic metrics. Furthermore, our analysis identifies specific hallucinations in generated summaries, including precedent citation errors and misrepresentations of case facts. These findings demonstrate the limitations of current automatic evaluation methods for legal summarization and highlight the critical role of human evaluation in assessing summary quality, particularly in complex, high-stakes domains. CaseSumm is available at https://huggingface.co/datasets/ChicagoHAI/CaseSumm
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。