用大模型和检索增强生成技术,高效提炼加尔各答高等法院判决并找相似案例
A Data Science Approach to Calcutta High Court Judgments: An Efficient LLM and RAG-powered Framework for Summarization and Similar Cases Retrieval
- 用微调的Pegasus模型进行两步摘要,保留关键法律上下文
- 构建向量数据库,支持快速检索相似案件,响应用户查询
- 适合法律研究者与学生,提升判例分析效率
司法作为民主三大支柱之一,正面临日益增长的法律案件压力,亟需高效利用司法资源。本研究提出一个融合数据科学方法的复杂框架,重点运用大型语言模型(LLM)与检索增强生成(RAG)技术,提升对加尔各答高等法院判决书的分析效率。该框架包含两项核心功能:一是构建鲁棒的摘要机制,将复杂的法律文本提炼为简洁连贯的摘要;二是开发智能相似案例检索系统,辅助法律从业者进行研究与决策。通过使用案件摘要头注对Pegasus模型进行微调,显著提升了法律案件摘要质量。采用两步摘要技术,有效保留了关键法律语境,生成全面的向量数据库以支持RAG。基于RAG的框架能高效响应用户查询,提供完整的相似案例概述与摘要。该方法不仅提高法律研究效率,也帮助法律专业人士与学生更便捷地获取和理解核心法律信息,改善整体法律环境。
原文摘要 · Abstract (English)
The judiciary, as one of democracy's three pillars, is dealing with a rising amount of legal issues, needing careful use of judicial resources. This research presents a complex framework that leverages Data Science methodologies, notably Large Language Models (LLM) and Retrieval-Augmented Generation (RAG) techniques, to improve the efficiency of analyzing Calcutta High Court verdicts. Our framework focuses on two key aspects: first, the creation of a robust summarization mechanism that distills complex legal texts into concise and coherent summaries; and second, the development of an intelligent system for retrieving similar cases, which will assist legal professionals in research and decision making. By fine-tuning the Pegasus model using case head note summaries, we achieve significant improvements in the summarization of legal cases. Our two-step summarizing technique preserves crucial legal contexts, allowing for the production of a comprehensive vector database for RAG. The RAG-powered framework efficiently retrieves similar cases in response to user queries, offering thorough overviews and summaries. This technique not only improves legal research efficiency, but it also helps legal professionals and students easily acquire and grasp key legal information, benefiting the overall legal scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。