arXiv:2510.17900cs.CYcs.AI2025-10被引 4

评测大模型在印度法律考试中的表现,发现其无法替代人类律师的深度推理。

Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning

  • 用印度真实法律考试构建评估基准,覆盖客观题与长文作答。
  • 大模型在客观题上接近或超过人类顶尖考生,但长文推理仍逊于人类第一名。
  • 模型在引用规范、格式合规和法律语境表达上存在明显缺陷,适合辅助而非替代。

大型语言模型正进入法律工作流程,但缺乏针对特定司法管辖区的评估框架。本文以印度公开法律考试为透明代理,构建了跨多年份的基准测试,整合了国家级和省级考试中的客观题,并在真实考试条件下评估开源及前沿大模型。为深入考察非选择题能力,还引入由律师评分、双盲对照的最高法院辩护人资格考试长文答卷研究。据我们所知,这是首个公开数据集与评估协议的、基于考试的印度法律领域大模型可法庭化评测体系。结果显示,尽管前沿模型在客观题上持续通过历史分数线,常达到甚至超越近年顶尖考生水平,但在长文推理任务中均未超越人类最高分者。评审意见一致指向三大可靠性失效模式:程序或格式合规性不足、权威引用不严谨、以及符合法庭语境的语体与结构缺失。这些发现明确了大模型可辅助的环节(如检查、跨法条一致性、法条与判例查询),也强调了人类主导不可或缺的领域:针对性文书起草与提交、程序与救济策略制定、权威冲突与例外协调,以及伦理与责任判断。

原文摘要 · Abstract (English)

Large language models (LLMs) are entering legal workflows, yet we lack a jurisdiction-specific framework to assess their baseline competence therein. We use India's public legal examinations as a transparent proxy. Our multi-year benchmark assembles objective screens from top national and state exams and evaluates open and frontier LLMs under real-world exam conditions. To probe beyond multiple-choice questions, we also include a lawyer-graded, paired-blinded study of long-form answers from the Supreme Court's Advocate-on-Record exam. This is, to our knowledge, the first exam-grounded, India-specific yardstick for LLM court-readiness released with datasets and protocols. Our work shows that while frontier systems consistently clear historical cutoffs and often match or exceed recent top-scorer bands on objective exams, none surpasses the human topper on long-form reasoning. Grader notes converge on three reliability failure modes: procedural or format compliance, authority or citation discipline, and forum-appropriate voice and structure. These findings delineate where LLMs can assist (checks, cross-statute consistency, statute and precedent lookups) and where human leadership remains essential: forum-specific drafting and filing, procedural and relief strategy, reconciling authorities and exceptions, and ethical, accountable judgment.

法律AI大模型评估印度法律推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。