剖析大模型在法律推理中的步骤错误,提出可量化的评估框架。
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
- 构建基于推理链的错误分类体系,识别逻辑漏洞
- 发现大模型在规则应用与类比推理中普遍出错
- 框架可复用于复杂逻辑任务的细粒度分析
近年来,大模型的推理能力备受关注。法律推理作为一项具有挑战性的领域,需精确应用规则与判例,平衡演绎与类比推理,并处理规则冲突。尽管已有研究尝试用大模型进行法律推理,但多聚焦于整体准确率。本文针对这一问题,深入分析大模型在法律推理中的逐步错误,使用来自《民事程序》数据集的大学水平多选题问答任务,基于对多个大模型推理链的初步人工分析,提出一种新的错误分类体系,并引入两个客观度量:合理性得分与正确性得分。进一步开发了基于大模型的自动化评估框架,用于识别推理错误并评估性能。通过该自动评估框架在数据集上计算合理性与正确性得分,揭示了若干有趣发现。此外,我们还证明将错误分类体系作为反馈融入主流提示技术,仅能小幅提升大模型表现。本工作可作为逻辑密集型复杂任务中推理链的详细错误分析评估框架。
原文摘要 · Abstract (English)
Reasoning abilities of LLMs have been a key focus in recent years. One challenging reasoning domain with interesting nuances is legal reasoning, which requires careful application of rules, and precedents while balancing deductive and analogical reasoning, and conflicts between rules. Although there have been a few works on using LLMs for legal reasoning, their focus has been on overall accuracy. In this paper, we dig deeper to do a step-by-step analysis and figure out where they commit errors. We use the college-level Multiple Choice Question-Answering (MCQA) task from the \textit{Civil Procedure} dataset and propose a new error taxonomy derived from initial manual analysis of reasoning chains with respect to several LLMs, including two objective measures: soundness and correctness scores. We then develop an LLM-based automated evaluation framework to identify reasoning errors and evaluate the performance of LLMs. The computation of soundness and correctness on the dataset using the auto-evaluator framework reveals several interesting insights. Furthermore, we show that incorporating the error taxonomy as feedback in popular prompting techniques marginally increases LLM performance. Our work will also serve as an evaluation framework that can be used in detailed error analysis of reasoning chains for logic-intensive complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。