arXiv:2608.29616cs.CL2026-08

JPO让法律判罚模型更懂逻辑,推理更符合真实判决结构。

JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

论文配图:JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction
图 1 · 摘自论文原文
  • 分四步构建标准化推理流程,用教师生成的解释监督训练。
  • 结合预测准确率、推理完整性和步骤一致性设计复合奖励函数。
  • 针对关键法律段落动态调整奖励,提升模型对重要推理环节的关注。

刑事判决预测需从案件事实中推断法条、罪名和量刑结果。不同于普通分类任务,该过程要求法条与事实匹配、罪名由法条支撑、量刑与罪名一致,形成结构化推理链条。现有方法仅优化最终标签,评估推理质量也依赖大模型生成的评分标准,反映的是模型内部偏好而非法律判决的内在逻辑。本文提出司法政策优化(JPO)框架,用于中文刑事案件判决预测的结构化法律推理。JPO首先利用教师模型生成的推理路径,监督标准化的四步推理流程;随后采用强化学习,基于判决预测质量、推理结构完整性及跨步骤一致性设计复合奖励。此外,引入逐标记优势重加权和自适应裁剪机制,增强对法律关键推理片段的关注。在多个开源语言模型及三个中文法律基准上的实验表明,相较于监督微调和强化学习基线,JPO能持续提升判决预测与推理质量。

原文摘要 · Abstract (English)

Criminal judgment prediction requires models to infer statutory articles, charges, and sentencing outcomes from case facts. Unlike standard classification tasks, it involves a structured reasoning process in which statutes should be matched with facts, charges should be justified by statutes, and sentencing outcomes should remain consistent with charges. Existing approaches optimize final labels, and while some have attempted to evaluate reasoning quality, their evaluations are indirect, often relying on LLM-generated rubrics that reflect model-internal preferences rather than the inherent logical structure of legal adjudication. We propose Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction. JPO first uses teacher-generated rationales to supervise a standardized four-step reasoning process, and then applies reinforcement learning with a composite reward over legal prediction quality, reasoning structure completeness, and cross-step consistency. JPO further introduces token-level advantage reweighting and adaptive clipping for legally salient reasoning segments. Experiments on multiple open-source language models and three Chinese legal benchmarks show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.

法律AI结构化推理强化学习刑事判决

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。