测试大模型能否识别法律引用是否真正支持论点,发现模型常因主题相关误判。
Is this Citation on Point?

- 通过替换案例或页码生成扰动引用,评估模型对论点支持的验证能力。
- 模型能识别93%-100%错误案例引用,但仅识别37%-83%错误页码引用。
- 模型易被主题相关性误导,需显式提示才能提升支持验证准确率。
2023年,纽约一名法官因律师使用ChatGPT生成幻觉引用而制裁两人。此类错误大多可通过数据库查询发现;更隐蔽的问题是引用真实案件但不支持所提论点——现有LLM法律应用评估普遍忽略此问题。本文通过在两个法律语料库中对真实引用进行受控扰动(替换案例或更改同一案件内的页码),研究论点级引用支持验证。评估了十四种模型配置。模型对错误案例的检测率达93%-100%,但对错误页码的检测率仅为37%-61%(法院意见)和52%-83%(法律备忘录)。当模型未能识别错误页码时,其依据主题相关性而非页面级支持接受引用。规模扩展与深度推理可缩小差距,但未完全解决:即使使用GPT-5.4高推理强度,仍分别漏检40%的法院意见与18%的备忘录中的页码错配。显式提示模型在引用页验证支持可提升召回率,但同时提高误报率。识别正确法律主题与验证引用支持是两种独立能力,当前模型常混淆二者。
原文摘要 · Abstract (English)
In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered -- a failure mode that existing evaluations of LLMs for legal use cases largely overlook. In this paper, we study proposition-level citation support verification through controlled perturbations of real legal citations obtained from two legal corpora, either replacing the cited case or changing only the pinpoint page within the same case. We evaluate fourteen model configurations on the resulting examples. Models catch 93-100% of wrong-case corruptions. They catch only 37-61% of wrong-pinpoint corruptions on court opinions and 52-83% on legal briefs. When models fail to catch wrong-pinpoint corruptions, they accept the citation based on topical overlap rather than page-level support. Scale and extended reasoning narrow the gap but do not close it: GPT-5.4 with high reasoning effort still misses 40% of pinpoint mismatches on court opinions and 18% on briefs. Prompting the model to verify support at the cited page improves recall, but it also raises the false positive rate. Recognizing the right legal topic and verifying support for the cited proposition are distinct capabilities, and current models conflate them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。