arXiv:2608.01292cs.CL2026-08

测试大模型在不同法域间法律推理的差异,发现模型难准确引用对应法源。

CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models

论文配图:CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 基于真实法律条文构建三法域同案异判题集
  • 6149个实例验证模型跨法域推理能力不足
  • 引入法源锚定评估指标,提升评测严谨性

法律推理具有明显的法域依赖性:相同事实可能在不同法律体系中适用不同规则并得出不同结论。现有基准极少评估大语言模型(LLMs)识别此类法域差异的能力,尤其当相同事实模式导致不同法律结果时。我们提出CrossLex,一个基于真实法律来源、针对中国、加州和德国三个法域的同案异判基准,涵盖合同、消费者、刑事、家庭和劳动法等55个法律议题。该基准构建了与法域对齐的问题、答案及支持引文,共包含6,149个实例,分属385个事实组,所有内容经法律专业人士审核。为分离基础法律知识与跨法域推理能力,定义三类任务:单一法域推理(T1)、联合跨法域比较(T2)和细粒度跨法域评估(T3)。我们还提出“法源锚定联合”(Grounded Joint)指标,联合评估答案正确性与法律来源依据性,并提供统一评测框架。大规模实验表明,尽管当前模型常能正确回答法律问题,但在跨法域引用法律条文方面表现不佳。我们希望CrossLex能推动未来面向法源锚定的跨法域法律推理研究。

原文摘要 · Abstract (English)

Legal reasoning is inherently jurisdiction-dependent: the same facts can call for different legal rules and yield different conclusions across legal systems. Yet existing benchmarks rarely evaluate whether large language models (LLMs) can recognize such jurisdiction-specific variation, especially when identical fact patterns lead to divergent legal outcomes.We introduce CrossLex, a same-fact, legal-source-grounded benchmark for evaluating cross-jurisdictional legal reasoning in LLMs across three jurisdictions: China, California, and Germany. Built from authoritative legal sources, CrossLex aligns 55 legal issues spanning contract, consumer, criminal, family, and labor law, and constructs jurisdiction-aligned questions paired with answers and supporting citations. In total, CrossLex contains 6,149 instances organized into 385 fact groups, with all legal issues, answers, and cited authorities reviewed by legal professionals.To disentangle basic legal knowledge from cross-jurisdictional reasoning, CrossLex defines three complementary tasks: single-jurisdiction reasoning (T1), joint cross-jurisdictional comparison (T2), and fine-grained cross-jurisdictional evaluation (T3). We further propose Grounded Joint, a metric that jointly assesses answer correctness and legal-source grounding, and provide a unified evaluation for streamlined benchmarking. Extensive experiments on representative LLMs show that, although current models can often answer legal questions correctly, they struggle to provide accurate cross-jurisdictional legal citations.We hope that CrossLex will facilitate future research on source-grounded cross-jurisdictional legal reasoning.

法律AI跨法域基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。