arXiv:2512.01020cs.AIcs.CL2025-12被引 3

用法律问题树评估大模型法律推理质量,发现覆盖与正确性缺一不可。

Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics

  • 将判决书转为法律问题树,作为评估推理轨迹的评分标准
  • 大模型推理能力受问题覆盖和正确性双重影响,二者缺一不可
  • 检索增强生成提升整体能力,强化学习提升正确性但降低覆盖

评估大模型在专业领域(如法律)生成的推理轨迹质量,对确保可信度与可解释性至关重要,但因任务本身复杂而极具挑战。本文提出 LEGIT(LEGal Issue Trees),一个大规模(24,000 实例)的专家级法律推理数据集,侧重于推理轨迹的评估。我们将法院判决转化为对立双方主张与法院结论的层级化问题树,作为评估推理轨迹问题覆盖与正确性的评分标准。通过人工专家标注及与粗糙评分标准的对比,验证了这些评分标准的可靠性。利用 LEGIT 数据集,我们发现:(1)大模型的法律推理能力严重受问题覆盖与正确性双重影响;(2)检索增强生成(RAG)与基于评分标准的强化学习(RL)带来互补收益,其中 RAG 提升整体推理能力,而 RL 提升正确性但导致覆盖下降。

原文摘要 · Abstract (English)

Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainability, yet remains challenging due to the inherent complexity of such reasoning tasks. We introduce LEGIT (LEGal Issue Trees), a novel large-scale (24K instances) expert-level legal reasoning dataset with an emphasis on reasoning trace evaluation. We convert court judgments into hierarchical trees of opposing parties' arguments and the court's conclusions, which serve as rubrics for evaluating the issue coverage and correctness of the reasoning traces. We verify the reliability of these rubrics via human expert annotations and comparison with coarse, less informative rubrics. Using the LEGIT dataset, we show that (1) LLMs' legal reasoning ability is seriously affected by both legal issue coverage and correctness, and that (2) retrieval-augmented generation (RAG) and RL with rubrics bring complementary benefits for legal reasoning abilities, where RAG improves overall reasoning capability, whereas RL improves correctness albeit with reduced coverage.

法律AI推理评估大模型测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。