arXiv:2601.03144cs.CLcs.AI2026-01

让AI通过日本律师考试,靠的是自我验证而非复杂拆解。

Self-Verification is All You Need To Pass The Japanese Bar Examination

  • 用真实考试格式训练模型,实现自我验证。
  • 首次在原题型下达到及格线以上得分。
  • 单模型自检比多智能体更有效,适合专业推理任务。

尽管大型语言模型(LLMs)进展迅速,但在高度专业化和结构化的考试中实现可靠表现仍是重大挑战。日本律师考试尤为严苛,不仅要求高级法律推理能力,还需严格遵循包含多项命题联合评分的复杂作答格式。虽然近期研究通过将问题分解为简单对错判断取得改进,但这些方法未在原始考试格式和评分标准下系统评估,无法确认其是否真正具备考试级能力。本文提出一种基于新构建数据集训练的自验证模型,该数据集完全复现了考试的真实格式与评分体系。模型在实际考试尺度下评估时,成绩超过官方及格线,据我们所知,这是首个无需改变原题结构或评分规则而通过日本律师考试的LLM。我们还对比了多智能体推理和基于分解的监督等策略,发现它们均无法达到同等性能。结果表明,格式忠实的监督与一致性验证至关重要,精心设计的单模型方法在高风险专业推理任务中可优于更复杂的系统。数据集与代码已公开。

原文摘要 · Abstract (English)

Despite rapid advances in large language models (LLMs), achieving reliable performance on highly professional and structured examinations remains a significant challenge. The Japanese bar examination is a particularly demanding benchmark, requiring not only advanced legal reasoning but also strict adherence to complex answer formats that involve joint evaluation of multiple propositions. While recent studies have reported improvements by decomposing such questions into simpler true--false judgments, these approaches have not been systematically evaluated under the original exam format and scoring scheme, leaving open the question of whether they truly capture exam-level competence. In this paper, we present a self-verification model trained on a newly constructed dataset that faithfully replicates the authentic format and evaluation scale of the exam. Our model is able to exceed the official passing score when evaluated on the actual exam scale, marking the first demonstration, to our knowledge, of an LLM passing the Japanese bar examination without altering its original question structure or scoring rules. We further conduct extensive comparisons with alternative strategies, including multi-agent inference and decomposition-based supervision, and find that these methods fail to achieve comparable performance. Our results highlight the importance of format-faithful supervision and consistency verification, and suggest that carefully designed single-model approaches can outperform more complex systems in high-stakes professional reasoning tasks. Our dataset and codes are publicly available.

法律AI自验证考试通过

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。