arXiv:2605.02011cs.CLcs.AI2026-05被引 6

用智能体+评分标准优化,让AI写判决书更准更合规

Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization

论文配图:Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization
图 1 · 摘自论文原文
  • 设计智能体动态检索法律条文和判例,提升信息召回
  • 引入评分引导优化,使生成内容符合司法逻辑与规范
  • 在JuDGE数据集上显著优于现有方法,准确率和质量双提升

自动化撰写判决书对提升司法效率至关重要,但面临法律信息全面检索与严谨逻辑推理的双重挑战。现有方法依赖标准检索增强生成和监督微调,常出现证据召回不足、虚构法条引用及逻辑错误。为此,我们提出Judge-R1统一框架,通过联合优化法律信息收集与判决书生成来解决该问题。首先,引入动态规划智能体从多源检索精确条文与判例;其次,采用基于组相对策略优化(GRPO)的强化学习阶段,结合综合法律奖励函数,确保符合司法标准与推理逻辑。在JuDGE基准上的大量实验表明,Judge-R1在法律准确性和生成质量上均显著优于当前最优基线。

原文摘要 · Abstract (English)

Automating the drafting of judgment documents is pivotal to judicial efficiency, yet it remains challenging due to the dual requirements of comprehensive retrieval of legal information and rigorous logical reasoning. Existing approaches, typically relying on standard Retrieval-Augmented Generation and Supervised Fine-Tuning, often suffer from insufficient evidence recall, hallucinated statutory references, and logically flawed legal reasoning. To bridge this gap, we propose Judge-R1, a unified framework designed to enhance LLM-based judgment document generation by jointly improving legal information collection and judgment document generation. First, we introduce Agentic Legal Information Collection, which employs a dynamic planning agent to retrieve precise statutes and precedents from multiple sources. Second, we implement Rubric-Guided Optimization, a reinforcement learning phase utilizing Group Relative Policy Optimization (GRPO) with a comprehensive legal reward function to enforce adherence to judicial standards and reasoning logic. Extensive experiments on the JuDGE benchmark demonstrate that Judge-R1 significantly outperforms state-of-the-art baselines in both legal accuracy and generation quality.

司法AI智能体判决书生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。