用合成数据和课程学习训练法律大模型,提升法律文本理解能力。
SynLexLM: Scaling Legal LLMs with Synthetic Data and Curriculum Learning
- 分阶段从简单到复杂法律文本训练,结合合成数据增强
- 在BigLaw-Bench和EUR-Lex-Sum上优于传统模型和微调版本
- 适合法律AI研究者与需要法律文本分析工具的开发者
大型语言模型虽强大,但在法律等专业领域常需大量微调和数据。通用预训练难以捕捉法律细节,且真实法律数据获取困难。本文提出SynLexLM,一种高效预训练法律大模型的新方法:采用课程学习策略,逐步从简单到复杂法律文本与问题推进,并利用Gemini Pro等模型生成合成问答对以缓解数据稀缺。目标是在BigLaw-Bench与EUR-Lex-Sum等法律基准上实现性能超越传统模型与微调版本。初步工作已生成反映法律推理的合成问答对,旨在提升法律文档分析与研究工具能力,推动高级法律AI的普及。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are powerful but often require extensive fine-tuning and large datasets for specialized domains like law. General-purpose pre-training may not capture legal nuances, and acquiring sufficient legal data is challenging. We introduce SynLexLM, a novel approach to efficiently pre-train a legal LLM. Our method employs curriculum learning, progressing from simple to complex legal texts and queries, combined with synthetic data augmentation using models like Gemini Pro to address data scarcity. We aim to achieve improved performance on legal benchmarks (BigLaw-Bench, EUR-Lex-Sum) compared to traditional models and fine-tuned versions. Preliminary work involves generating synthetic QA pairs reflecting legal reasoning. This work aims to enhance legal document analysis and research tools, potentially democratizing access to advanced legal AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。