构建印度最大法律判例数据集,训练出可解释的法律AI模型。
NyayaAnumana & INLegalLlama: The Largest Indian Legal Judgment Prediction Dataset and Specialized Language Model for Enhanced Decision Analysis
- 整合70万+印度各级法院案件,构建跨语言多源法律数据集
- 基于大模型微调,实现90%判例预测准确率并生成可解释推理
- 适合法律AI研究者与司法智能化推进者使用
人工智能在法律判例预测(LJP)中的应用有望改变法律体系,尤其在案件积压严重的印度。本文提出NyayaAnumana,目前最大且最多样化的印度判例数据集,包含702,945条预处理后的案件,覆盖最高法院、高等法院、特别法庭、地方法院及日常裁决,涵盖主要印度语种。该数据集超越现有PredEx和ILDC等数据集,为法律领域先进AI研究提供坚实基础。同时,我们推出INLegalLlama——一种针对印度法律体系定制的生成式大语言模型,采用两阶段训练:先通过持续预训练注入印度法律文档,再进行任务特定监督微调。实验表明,融合多样化法院数据显著提升模型性能,在预测任务中达到约90%的F1分数。INLegalLlama不仅提高预测精度,还提供可理解的解释,满足法律决策中对可解释性的需求。
原文摘要 · Abstract (English)
The integration of artificial intelligence (AI) in legal judgment prediction (LJP) has the potential to transform the legal landscape, particularly in jurisdictions like India, where a significant backlog of cases burdens the legal system. This paper introduces NyayaAnumana, the largest and most diverse corpus of Indian legal cases compiled for LJP, encompassing a total of 7,02,945 preprocessed cases. NyayaAnumana, which combines the words "Nyay" (judgment) and "Anuman" (prediction or inference) respectively for most major Indian languages, includes a wide range of cases from the Supreme Court, High Courts, Tribunal Courts, District Courts, and Daily Orders and, thus, provides unparalleled diversity and coverage. Our dataset surpasses existing datasets like PredEx and ILDC, offering a comprehensive foundation for advanced AI research in the legal domain. In addition to the dataset, we present INLegalLlama, a domain-specific generative large language model (LLM) tailored to the intricacies of the Indian legal system. It is developed through a two-phase training approach over a base LLaMa model. First, Indian legal documents are injected using continual pretraining. Second, task-specific supervised finetuning is done. This method allows the model to achieve a deeper understanding of legal contexts. Our experiments demonstrate that incorporating diverse court data significantly boosts model accuracy, achieving approximately 90% F1-score in prediction tasks. INLegalLlama not only improves prediction accuracy but also offers comprehensible explanations, addressing the need for explainability in AI-assisted legal decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。