测试大模型对判例推翻关系的理解能力,发现其存在时间偏见和浅层推理缺陷。
Do LLMs Truly Understand When a Precedent Is Overruled?
- 构建236个美国最高法院案例对,评估大模型识别判例推翻关系的能力。
- 模型在历史案件上表现显著下降,且常生成时间逻辑错误的推论。
- 揭示大模型缺乏深层法律理解,适合关注法律AI真实能力的研究者参考。
大型语言模型(LLMs)在长上下文场景下展现出解决复杂法律推理任务的潜力,但其理解长篇法律文件的能力仍缺乏充分评估。现有评价多依赖简化合成任务,难以反映真实世界文档理解的复杂性。判例推翻关系是普通法体系的核心,常见于司法意见中,是检验长文档法律理解的理想场景。本文基于236个美国最高法院案例对,评估顶尖大模型识别判例推翻关系的表现。结果显示三大局限:(1) 时代敏感性——模型在历史案件上性能明显下降,暴露训练数据中的时间偏见;(2) 浅层推理——模型依赖表面逻辑启发式,而非深层法律理解;(3) 上下文依赖推理失败——在复杂开放式任务中生成时间上不可能的关系,尽管在简单情境中仍保持基本时间意识。本研究填补了真实长上下文评估的空白,提供一个贴近实际法律推理复杂性与高风险性的评测环境。
原文摘要 · Abstract (English)
Large language models (LLMs) with extended context windows show promise for complex legal reasoning tasks, yet their ability to understand long legal documents remains insufficiently evaluated. Developing long-context benchmarks that capture realistic, high-stakes tasks remains a significant challenge in the field, as most existing evaluations rely on simplified synthetic tasks that fail to represent the complexity of real-world document understanding. Overruling relationships are foundational to common-law doctrine and commonly found in judicial opinions. They provide a focused and important testbed for long-document legal understanding that closely resembles what legal professionals actually do. We present an assessment of state-of-the-art LLMs on identifying overruling relationships from U.S. Supreme Court cases using a dataset of 236 case pairs. Our evaluation reveals three critical limitations: (1) era sensitivity -- the models show degraded performance on historical cases compared to modern ones, revealing fundamental temporal bias in their training; (2) shallow reasoning -- models rely on shallow logical heuristics rather than deep legal comprehension; and (3) context-dependent reasoning failures -- models produce temporally impossible relationships in complex open-ended tasks despite maintaining basic temporal awareness in simple contexts. Our work contributes a benchmark that addresses the critical gap in realistic long-context evaluation, providing an environment that mirrors the complexity and stakes of actual legal reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。