arXiv:2604.23730cs.AI2026-04中稿 · ICAIL 2026

首个评估大模型日本法律推理能力的数据集,专家实测表现短板。

Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task

  • 基于日本司法考试写作题构建数据集,模拟真实法律场景
  • 专家评估发现模型在法律论证结构与事实关联上存在明显缺陷
  • 揭示模型幻觉模式,适合法律AI研究者与司法科技从业者参考

大型语言模型(LLMs)在法律基准测试中表现出色,尤其在选择题部分。然而,其在真实情境下生成开放式法律推理的能力仍缺乏充分研究。值得注意的是,目前尚无针对日本法律体系的此类研究或数据集。本研究首次构建了用于评估大模型在日本法域内开放性法律推理能力的数据集,其内容源自日本司法考试的写作部分,要求考生从长篇叙述中识别多个法律问题,并以自由文本形式构建结构化法律论证。核心贡献在于由法律专家对模型生成的回答进行人工评估,揭示了模型在法律推理上的局限与挑战。此外,我们还对幻觉现象进行了人工分析,刻画出模型在何种情况下引入未经判例或法律支持的内容。真实考题、模型输出及专家评价共同展现了当前大模型在日语法律领域的发展水平。相关数据集与资源将公开发布。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown strong performance on legal benchmarks, including multiple-choice components of bar exams. However, their capacity for generating open-ended legal reasoning in realistic scenarios remains insufficiently explored. Notably, to our best knowledge, there are no prior studies or datasets addressing this issue in the Japanese context. This study presents the first dataset designed to evaluate the open-ended legal reasoning performance of LLMs within the Japanese jurisdiction. The dataset is based on the writing component of the Japanese bar examination, which requires examinees to identify multiple legal issues from long narratives and to construct structured legal arguments in free text format. Our key contribution is the manual evaluation of LLMs' generated responses by legal experts, which reveals limitations and challenges in legal reasoning. Moreover, we conducted a manual analysis of hallucinations to characterize when and how the models introduce content not supported by precedent or law. Our real exam questions, model-generated responses, and expert evaluations reveal the milestones of current LLMs in the Japanese legal domain. Our dataset and relevant resources will be available online.

法律AI大模型评测日本法幻觉分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。