arXiv:2606.00898cs.CLcs.DL2026-06

法律大模型幻觉率评估依赖引用图覆盖率,而非模型本身表现。

Citation Grounding Measures the Oracle: Graph Coverage Determines Reported LLM Hallucination Rates in Law

  • 用真实判决书构建引用图作为验证基准,不依赖人工标注。
  • 图覆盖稀疏时幻觉率显示15%-21%,密集时降至0.1%以下,结果差异源于覆盖率。
  • 高覆盖率图虽可信但无法区分模型优劣,低覆盖率图却误判采集质量。

将LLM生成的法律引文与从真实判决书中提取的引用图进行比对,是大规模衡量幻觉的理想方式:无需标注、无需参考答案。我们发现,该度量报告的结果由所查询图的覆盖率决定,而非被评估模型。当覆盖率足够高使结论可信时,该指标已完全无法区分不同模型。对400条响应(100个乌克兰法律问题×4个商用大模型)在两个时间点的同一国家引用图快照上评分:稀疏快照(4.7e5条记录)下引用准确率0.791-0.855,显示15%-21%引文为幻觉;密集快照(3.3e8条记录,5.8e7份判决)下得分0.989-0.999。子采样实验表明,差异源于覆盖率——模拟均匀采样仅凭记录数即可复现稀疏结果,误差小于0.018,且不依赖模型信息。对查询的自助抽样显示,任何规模的基准下,无一对系统在95%置信水平上可区分。该度量陷入双重失败:稀疏基准虽能区分,但区分的是采样覆盖;密集基准可信,却无法区分模型。提升图密度虽解决第一问题,却引发第二问题。独立立法注册表验证:稀疏基准标记的54条引文全部真实,假阳性率100%;密集基准标记的4条中,2条为虚构,1条是覆盖盲区,1条无法分类——疑似废止条款,构成论证缩影。

原文摘要 · Abstract (English)

Verifying LLM-generated legal citations against a graph of citations extracted from real court decisions is an appealing way to measure hallucination at scale: no annotators, no reference answers. We show that what such a metric reports is governed by the coverage of the graph it queries rather than by the model it evaluates, and that at the coverage where its verdicts become trustworthy it stops distinguishing models at all. We score 400 responses (100 Ukrainian legal queries x four commercial LLMs) against two snapshots of the same national citation graph, holding responses, extractor and metric fixed. Against a sparse snapshot (4.7e5 records) citation grounding ranges 0.791-0.855, apparently showing 15-21% of citations hallucinated. Against a dense snapshot of the same registry re-derived ten weeks later (3.3e8 records, 5.8e7 decisions) the identical responses score 0.989-0.999. Subsampling shows the cause is coverage: harvesting modelled as uniform record sampling, calibrated on nothing but the record count, reproduces the sparse scores to within 0.018 while knowing nothing about the models. Bootstrapping over the queries shows the other half, which no version of this work reported: no pair of systems is separable at 95% at any oracle size tested. The metric is caught between two failures. Sparse oracles discriminate, but what they discriminate is harvesting coverage; dense oracles are trustworthy and separate nothing. Densifying the graph, the obvious remedy for the first, produces the second. An independent legislation registry adjudicates: all 54 citations flagged by the sparse oracle name real articles, a 100% false-positive rate; of the four flagged by the dense oracle, two are fabrications, one a coverage gap, and one we cannot classify - it looks like a repealed provision, which is the argument in miniature.

法律AI幻觉评估引用图评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。