测试大模型在法律推理中的分层分析能力,发现越思考越容易出错。
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
- 将法律案例拆解为三阶段推理任务,用因素和知识层级建模
- 高层级推理准确率仅64.82%~92.09%,综合分析低至11.46%~33.99%
- 计算资源多用在错误答案上,提示‘更长时间思考’未必更聪明
判例推理是美国法律实践的核心,要求从业者通过类比与区分过往判例来论证当前案件。尽管大语言模型(LLMs)表现卓越,其在这一复杂、细微的推理形式中的能力仍需深入探究。我们提出一个形式化框架,将识别案例间关键差异的过程分解为三个阶段的推理任务:使用称为‘因素’的事实谓词建模案例,构建法律知识层次结构,并定义可验证规则以识别差异、分析其论据支持并评估其重要性。通过对现代推理型大模型的全面评估,我们揭示了一个悖论:模型在表层推理(任务1)中表现良好,但在层级推理(任务2:准确率64.82%–92.09%)中性能下降,在综合分析(任务3:准确率11.46%–33.99%)中几乎崩溃。最显著的是,模型在错误回答上消耗的计算资源远高于正确回答,表明‘思考更久’并不等于‘思考更聪明’。本研究提供了一种对复杂领域中大模型推理能力进行细粒度分析的方法,揭示了法律人工智能实现稳健可信所必须克服的根本局限。
原文摘要 · Abstract (English)
Case-based reasoning is a cornerstone of U.S. legal practice, requiring professionals to argue about a current case by drawing analogies to and distinguishing from past precedents. While Large Language Models (LLMs) have shown remarkable capabilities, their proficiency in this complex, nuanced form of reasoning needs further investigation. We propose a formal framework that decomposes the process of identifying significant distinctions between cases into three-stage reasoning tasks. Our framework models cases using factual predicates called factors, organizes them into a legal knowledge hierarchy, and defines verifiable rules for identifying distinctions, analyzing their argumentative support, and evaluating their significance. Through comprehensive evaluation of modern reasoning LLMs, we reveal a paradox: while models achieve high accuracy on surface-level reasoning (Task 1), performance degrades on hierarchical reasoning (Task 2: 64.82%-92.09%) and collapses on integrated analysis (Task 3: 11.46%-33.99%). Most strikingly, we find that models consistently expend more computational resources on incorrect responses than correct ones, suggesting that "thinking longer" does not always mean "thinking smarter." Our work provides a methodology for fine-grained analysis of LLM reasoning capabilities in complex domains and reveals fundamental limitations that must be addressed for robust and trustworthy legal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。