在真实庭审场景下测试大模型判案能力,发现GPT-3.5表现最优但仍未达专家水平。
Rethinking Legal Judgement Prediction in a Realistic Scenario in the Era of Large Language Models
- 模拟法庭实时决策,仅用案情、法条和判例作为输入
- GPT-3.5 Turbo在真实场景中表现最佳,准确率显著优于其他模型
- 提供可解释预测,但人类评估显示仍不及法律专家
本研究在印度司法背景下,针对真实庭审场景下的判决预测问题展开探究,采用InLegalBERT、BERT、XLNet等基于Transformer的模型以及Llama-2、GPT-3.5 Turbo等大语言模型(LLMs)。实验模拟案件提交法院时的决策时刻,仅使用当时可获取的信息,如案件事实、法律条文、判例及诉讼主张,避免事后分析的偏差。针对Transformer模型,尝试分层结构与事实摘要优化输入。实验表明,GPT-3.5 Turbo在真实场景中表现突出,且引入法条与判例能显著提升预测效果。此外,LLMs可生成预测解释。为评估预测与解释质量,提出两个人工评价指标:清晰度与关联性。自动与人工评估均显示,尽管大模型进展显著,其判决预测与解释能力仍未能达到法律专家水平。
原文摘要 · Abstract (English)
This study investigates judgment prediction in a realistic scenario within the context of Indian judgments, utilizing a range of transformer-based models, including InLegalBERT, BERT, and XLNet, alongside LLMs such as Llama-2 and GPT-3.5 Turbo. In this realistic scenario, we simulate how judgments are predicted at the point when a case is presented for a decision in court, using only the information available at that time, such as the facts of the case, statutes, precedents, and arguments. This approach mimics real-world conditions, where decisions must be made without the benefit of hindsight, unlike retrospective analyses often found in previous studies. For transformer models, we experiment with hierarchical transformers and the summarization of judgment facts to optimize input for these models. Our experiments with LLMs reveal that GPT-3.5 Turbo excels in realistic scenarios, demonstrating robust performance in judgment prediction. Furthermore, incorporating additional legal information, such as statutes and precedents, significantly improves the outcome of the prediction task. The LLMs also provide explanations for their predictions. To evaluate the quality of these predictions and explanations, we introduce two human evaluation metrics: Clarity and Linking. Our findings from both automatic and human evaluations indicate that, despite advancements in LLMs, they are yet to achieve expert-level performance in judgment prediction and explanation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。