arXiv:2505.01539cs.AIcs.LG2025-05中稿 · presentation as a …被引 2

用可调复杂度的论证推理题评估大模型法律推理能力

Parameterized Argumentation-based Reasoning Tasks for Benchmarking Generative Language Models

  • 基于证人证词构建动态可变的论证攻击图,生成自然语言推理题
  • 主流大模型在低复杂度下即出错,高复杂度时连专为推理设计的模型也失败
  • 适合评估法律AI系统可靠性,助力负责任AI研发

生成式大语言模型在法律领域有提升司法系统的潜力,但其推理行为脆弱且难以理解,难以在法律与证据领域负责任地应用。本文提出一种创建基准测试的方法,用于评估生成式语言模型的推理能力。这些基准具有动态变化、可扩展复杂度和形式上无歧义的特点。研究以证人证词为基础,聚焦论证攻击结构,动态生成线性和非线性论证攻击图,涵盖不同复杂度,并将其转化为自然语言表达的推理谜题。实验表明,当前最先进大模型在低复杂度下便频繁出错,表现不一致,显示其推理能力脆弱;在更高复杂度下,即使专为推理优化的模型也出现错误。结果证明,参数化复杂度的基准可用于有效评估生成式模型的推理能力,有助于更深入理解其局限性,对设计法律领域负责任的AI系统至关重要。

原文摘要 · Abstract (English)

Generative large language models as tools in the legal domain have the potential to improve the justice system. However, the reasoning behavior of current generative models is brittle and poorly understood, hence cannot be responsibly applied in the domains of law and evidence. In this paper, we introduce an approach for creating benchmarks that can be used to evaluate the reasoning capabilities of generative language models. These benchmarks are dynamically varied, scalable in their complexity, and have formally unambiguous interpretations. In this study, we illustrate the approach on the basis of witness testimony, focusing on the underlying argument attack structure. We dynamically generate both linear and non-linear argument attack graphs of varying complexity and translate these into reasoning puzzles about witness testimony expressed in natural language. We show that state-of-the-art large language models often fail in these reasoning puzzles, already at low complexity. Obvious mistakes are made by the models, and their inconsistent performance indicates that their reasoning capabilities are brittle. Furthermore, at higher complexity, even state-of-the-art models specifically presented for reasoning capabilities make mistakes. We show the viability of using a parametrized benchmark with varying complexity to evaluate the reasoning capabilities of generative language models. As such, the findings contribute to a better understanding of the limitations of the reasoning capabilities of generative models, which is essential when designing responsible AI systems in the legal domain.

法律AI推理评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。