arXiv:2604.10034cs.AI2026-04

语言模型首次零错误完成官方LSAT考试,证明其法律推理能力已达人类顶尖水平。

AI Achieves a Perfect LSAT Score

  • 通过分析思维过程对模型表现影响,发现推理阶段是高分关键。
  • 经优化后模型在逻辑推理题上实现满分,较蒸馏模型提升显著。
  • 适合关注AI认知极限、法律推理与大模型可解释性的研究者阅读。

本文首次记录了语言模型在官方公布的法学院入学考试(LSAT)中取得满分的成绩。对八种推理模型的受控实验表明,提示词变化、选项顺序打乱及多次采样均未显著影响性能。然而,移除模型生成的答案前思考阶段,使前沿模型准确率下降最多达8个百分点,主要体现在逻辑推理题上。蒸馏模型虽能生成完整思考过程,但表现远低于前沿水平。通过在官方LSAT解析数据上使用QLoRA微调的奖励模型进行Best-of-5选择,缩小了这一差距,提升仍主要集中于逻辑推理部分。自1948年以来作为精英法律教育准入门槛的LSAT,如今已被能够无误作答的语言模型突破。它所测试的认知能力上限,已不再专属于人类认知。

原文摘要 · Abstract (English)

This paper reports the first documented instance of a language model achieving a perfect score on an officially disclosed Law School Admission Test (LSAT). Controlled experiments on eight reasoning models show that varying the prompt, shuffling answer choices, and sampling multiple responses have no meaningful effect as drivers of performance. Ablating the thinking phase that models generate before answering, however, lowers frontier accuracy by up to 8 percentage points, predominantly in logical reasoning. Distilled models produce full thinking traces in the same format yet plateau far below frontier performance. A pilot process reward model fine-tuned via QLoRA on official LSAT explanations narrows this gap through Best-of-5 selection, with gains again predominantly in logical reasoning. The gatekeeper of elite legal education since 1948, the LSAT has not merely been passed but answered without a single error by models that reason. The upper bound of the cognitive capacities it has tested is no longer exclusive to human cognition.

语言模型法律推理认知能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。