法律大模型答对题却漏引法条,现有评测标准会误判为正确。
Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks
- 在未要求引用法条的提示下,4个大模型在238道台湾司法考试题中自发标注法条。
- 24.0%至42.4%的正确答案未引用正确法条,15.2%至21.7%的错误答案却引用了正确法条。
- 提出联合评估答案与法条引用的新标准,适合法律AI评测与跨法域研究者。
法律基准测试通常仅评分最终答案,即使模型同时陈述法律依据。我们检验答案正确性能否作为依据扎根的代理指标。在未要求引用法条的常规推理提示下,四个大语言模型在238道台湾司法考试题中自发产生法条标记。因每道题均有已验证的管辖条款,可自动联合审计答案正确性与法条引用准确性。两者在两个方向上均出现分离:在刑法领域,24.0–42.4%的正确回答未引用黄金法条,15.2–21.7%的错误回答却引用了它。独立的法条检索探测与允许不引用的干预实验进一步表明,答案与引用行为可在输出层面独立变化。由于这种错配在非对抗性提示下自然发生,仅以答案正确性评分会将本应识别的黄金法条遗漏误判为完整成功。鉴于法条具有结构性可提取性与外部可验证性,该失败可被自动测量。一项初步的中国民法扩展也观察到未请求引用时的法条标记现象,推动开展全跨法域联合审计。因此,我们建议对以法条为基础的法律基准测试采用联合答案-法条评估。
原文摘要 · Abstract (English)
Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items. Because each item has a verified governing provision, we automatically audit answer correctness and authority grounding jointly. The two dimensions dissociate in both directions. In criminal law, 24.0--42.4\% of valid responses were answer-correct but missed the gold authority, while 15.2--21.7\% were answer-incorrect but cited it. A separate statutory-retrieval probe and a permissive citation-abstention intervention further show that answer and citation behavior can move separately at the output level. Because this mismatch arises without adversarial or inconsistency-inducing prompting, answer-only scoring treats naturally occurring gold-authority misses as complete benchmark successes. Because statutory authority is structurally extractable and externally verifiable, the failure can be measured automatically. A preliminary PRC civil-law extension also observes citation-unrequested authority marking, motivating a full cross-jurisdictional joint audit. We therefore propose joint answer--authority evaluation for statute-grounded legal benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。