arXiv:2602.18346cs.CLcs.AI2026-02

用AI预测印度上诉判决并生成可解释的法律推理,缓解司法积案压力。

Vichara: Appellate Judgment Prediction and Explanation for the Indian Judicial System

  • 将上诉案件拆解为法律要点,结构化提取核心判断与上下文。
  • 在两个数据集上达到最高81.5的F1分数,GPT-4o mini表现最佳。
  • 解释符合印度法理逻辑,适合法官、律师等专业人员使用。

在印度等司法体系面临严重案件积压的背景下,人工智能为法律判决预测带来变革可能。本文聚焦上诉案件——即高等法院对下级法院裁决的正式复审,提出专为印度司法体系设计的Vichara框架,实现上诉判决的预测与解释。该框架处理英文上诉案件文书,将其分解为决策点(decision points),每个点包含法律问题、裁决主体、结果、推理及时间背景。这种结构化表示剥离核心法律判断及其上下文,支持精准预测与可解释性分析。解释采用基于IRAC框架但适配印度法律推理的格式,提升可读性,使法律从业者能高效评估预测合理性。我们在PredEx和专家标注的印度法律文档语料库子集(ILDC_expert)上评估了四种大语言模型:GPT-4o mini、Llama-3.1-8B、Mistral-7B和Qwen2.5-7B。Vichara在两数据集上均超越现有基准,其中GPT-4o mini表现最优(PredEx F1: 81.5,ILDC_expert F1: 80.3),其次为Llama-3.1-8B。人类评估显示其解释在清晰度、关联性和实用性方面均具优势,尤其以GPT-4o mini生成内容最佳。

原文摘要 · Abstract (English)

In jurisdictions like India, where courts face an extensive backlog of cases, artificial intelligence offers transformative potential for legal judgment prediction. A critical subset of this backlog comprises appellate cases, which are formal decisions issued by higher courts reviewing the rulings of lower courts. To this end, we present Vichara, a novel framework tailored to the Indian judicial system that predicts and explains appellate judgments. Vichara processes English-language appellate case proceeding documents and decomposes them into decision points. Decision points are discrete legal determinations that encapsulate the legal issue, deciding authority, outcome, reasoning, and temporal context. The structured representation isolates the core determinations and their context, enabling accurate predictions and interpretable explanations. Vichara's explanations follow a structured format inspired by the IRAC (Issue-Rule-Application-Conclusion) framework and adapted for Indian legal reasoning. This enhances interpretability, allowing legal professionals to assess the soundness of predictions efficiently. We evaluate Vichara on two datasets, PredEx and the expert-annotated subset of the Indian Legal Documents Corpus (ILDC_expert), using four large language models: GPT-4o mini, Llama-3.1-8B, Mistral-7B, and Qwen2.5-7B. Vichara surpasses existing judgment prediction benchmarks on both datasets, with GPT-4o mini achieving the highest performance (F1: 81.5 on PredEx, 80.3 on ILDC_expert), followed by Llama-3.1-8B. Human evaluation of the generated explanations across Clarity, Linking, and Usefulness metrics highlights GPT-4o mini's superior interpretability.

司法AI判决预测可解释性法律推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。