注入法律领域知识提升印式判決摘要效果,支持英印双语输出。
Advantages of Domain Knowledge Injection for Legal Document Summarization: A Case Study on Summarizing Indian Court Judgments in English and Hindi
- 用法律专有编码器增强抽取式模型,融合领域知识。
- 通过持续预训练使生成模型在英印双语法律语料上表现更优。
- 专家验证表明方法有效,适合法律AI与多语言应用者。
印度法律判决摘要是一项复杂任务,不仅因法律文本语言繁复、结构不规整,还因大量民众无法理解英文法律文件,需提供印地语等本地语言摘要。本研究旨在通过注入领域知识,提升英、印双语法律文本摘要性能。提出框架:将法律专用预训练编码器融入抽取式神经模型;并通过大规模英、印双语法律语料对生成模型(含大语言模型)进行持续预训练。实验显示,该方法在标准评估指标、事实一致性及法律领域特有指标上均达统计显著提升。专家评审进一步验证了方法有效性。
原文摘要 · Abstract (English)
Summarizing Indian legal court judgments is a complex task not only due to the intricate language and unstructured nature of the legal texts, but also since a large section of the Indian population does not understand the complex English in which legal text is written, thus requiring summaries in Indian languages. In this study, we aim to improve the summarization of Indian legal text to generate summaries in both English and Hindi (the most widely spoken Indian language), by injecting domain knowledge into diverse summarization models. We propose a framework to enhance extractive neural summarization models by incorporating domain-specific pre-trained encoders tailored for legal texts. Further, we explore the injection of legal domain knowledge into generative models (including Large Language Models) through continual pre-training on large legal corpora in English and Hindi. Our proposed approaches achieve statistically significant improvements in both English-to-English and English-to-Hindi Indian legal document summarization, as measured by standard evaluation metrics, factual consistency metrics, and legal domain-specific metrics. Furthermore, these improvements are validated through domain experts, demonstrating the effectiveness of our approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。