arXiv:2511.22869cs.CL2025-11被引 1

构建首个日语法律领域大模型评测数据集,覆盖三大法典。

JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge

  • 从日本律师考试多选题中提取3464道题,拆分为独立判断项
  • 宪法题普遍比民法、刑法题更易答对,闭源推理模型表现最优
  • 适合评估日语法律大模型能力,尤其关注逻辑推理与法律知识

我们提出JBE-QA,一个用于评估大语言模型法律知识的日语法律问答数据集。该数据集源自2015至2024年日本律师考试的多选题部分,是首个针对日语法律领域的大规模基准评测数据。涵盖民法典、刑法典和宪法,突破了以往资源集中于民法典的局限。每道题目被分解为独立的真/假判断,并配有结构化上下文字段。数据集共包含3,464个样本,标签分布均衡。我们评估了26个大语言模型,包括闭源模型、开源权重模型、日语专用模型及具备推理能力的模型。结果表明,开启推理功能的闭源模型表现最佳,且宪法相关问题整体上比民法或刑法问题更易作答。

原文摘要 · Abstract (English)

We introduce JBE-QA, a Japanese Bar Exam Question-Answering dataset to evaluate large language models' legal knowledge. Derived from the multiple-choice (tanto-shiki) section of the Japanese bar exam (2015-2024), JBE-QA provides the first comprehensive benchmark for Japanese legal-domain evaluation of LLMs. It covers the Civil Code, the Penal Code, and the Constitution, extending beyond the Civil Code focus of prior Japanese resources. Each question is decomposed into independent true/false judgments with structured contextual fields. The dataset contains 3,464 items with balanced labels. We evaluate 26 LLMs, including proprietary, open-weight, Japanese-specialised, and reasoning models. Our results show that proprietary models with reasoning enabled perform best, and the Constitution questions are generally easier than the Civil Code or the Penal Code questions.

法律AI日语模型评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。