arXiv:2605.10186cs.CLcs.AI2026-05被引 1

测试法律大模型在无外部资料时引用真实判例的能力,发现多数模型会自信地编造错误案例。

LegalCiteBench: Evaluating Citation Reliability in Legal Language Models

论文配图:LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
图 1 · 摘自论文原文
  • 构建24000个真实美国判例任务,评估模型闭卷检索与验证引用能力
  • 最强模型引用准确率不足7%,94%以上模型在检索中给出误导性答案
  • 即使增加训练数据或模型规模,仍难解决编造引用问题,适合研究法律AI可靠性

大型语言模型正被广泛用于法律起草与研究,但错误引用或虚构判例可能带来严重专业后果。现有法律评测多聚焦法规推理、合同理解或通用法律问答,却未直接考察普通法中的核心缺陷:当无法依赖外部资料时,模型可能生成看似合理实则错误的引用或判例。本文提出LegalCiteBench,一个用于评估法律大模型在闭卷条件下进行引用恢复、验证与判例匹配的基准。该基准基于1000份真实的美国司法判例(来自Case Law Access Project),包含约24,000个评估实例,涵盖五项以引用为核心的任务:引用检索、引用补全、引用错误检测、判例匹配与判例验证修正。在21个大模型中,即使最强模型在引用检索与补全任务上的精确率也低于7/100。模型规模和法律领域预训练带来的提升有限,无法解决此难题。在评估协议下,多数模型频繁提供具体但错误或重叠度低的判例,其中20个模型在检索密集型任务上的误导性回答率(MAR)超过94%。仅通过提示引导放弃回答的实验表明,明确表达不确定性可减少部分自信造假,但未能提升引用准确性。LegalCiteBench旨在作为诊断框架,用于研究模型在缺乏外部支撑时的权威生成失败、验证行为及回避策略。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly integrated into legal drafting and research workflows, where incorrect citations or fabricated precedents can cause serious professional harm. Existing legal benchmarks largely emphasize statutory reasoning, contract understanding, or general legal question answering, but they do not directly study a central common-law failure mode: when asked to provide case authorities without external grounding, models may return plausible-looking but incorrect citations or cases. We introduce LegalCiteBench, a benchmark for studying closed-book citation recovery, citation verification, and case matching in legal language models. LegalCiteBench contains approximately 24K evaluation instances constructed from 1,000 real U.S. judicial opinions from the Case Law Access Project. The benchmark covers five citation-centric tasks: citation retrieval, citation completion, citation error detection, case matching, and case verification and correction. Across 21 LLMs, exact citation recovery remains highly challenging in this closed-book setting: even the strongest models score below 7/100 on citation retrieval and completion. Within the evaluated models, scale and legal-domain pretraining provide limited gains and do not resolve this difficulty. Models also frequently provide concrete but incorrect or low-overlap authorities under our evaluation protocol, with Misleading Answer Rates (MAR) exceeding 94% for 20 of 21 evaluated models on retrieval-heavy tasks. A prompt-only abstention experiment shows that explicit uncertainty instructions reduce some confident fabrication but do not improve citation correctness. LegalCiteBench is intended as a diagnostic framework for studying authority generation failures, verification behavior, and abstention when external grounding is absent, incomplete, or bypassed.

法律AI引用可靠性闭卷评测模型幻觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。