arXiv:2410.21479cs.CLcs.AI2024-10被引 4

用大模型自动生成法律阅读理解数据,低成本打造专业法律大模型。

TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text

  • 用大模型将原始法律文本转为阅读理解题,提升训练效率。
  • 在法律基准测试中超越更大规模模型,仅用5亿+法律token继续预训练。
  • 方法可推广至其他领域,适合资源有限但需专业模型的团队。

大型语言模型(LLMs)在高度专业化领域展现出潜力,但在准确性和成本方面仍存挑战,限制了其在特定任务中的应用。尽管微调预训练模型已取得良好效果,但该过程计算开销大,且需要大量专有数据。本文提出新方法:基于Phi-2和Mistral-7B-v0.1,使用超过5亿个法律文本令牌进行持续预训练,构建了针对法律场景的Phi-2-Legal与Mistral-Legal-7B模型。通过利用大模型将原始训练数据转化为阅读理解形式,显著提升了法律任务能力。实验表明,这些模型在法律基准测试中表现优异,甚至超越了使用更大数据集和更多资源训练的模型。本工作强调了领域特定文本持续预训练的有效性,并展示了以低成本大模型完成数据转换的可行性,使模型兼具领域专长与通用语言理解能力。该方法不仅适用于法律领域,还可扩展至任意预训练数据集,推动各类任务性能提升。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown promise in highly-specialized domains, however challenges are still present in aspects of accuracy and costs. These limitations restrict the usage of existing models in domain-specific tasks. While fine-tuning pre-trained models have shown promising results, this process can be computationally expensive and require massive datasets of the specialized application in hand. In this work, we bridge that gap. We have developed Phi-2-Legal and Mistral-Legal-7B, which are language models specifically designed for legal applications. These models are based on Phi-2 and Mistral-7B-v0.1, and have gone through continued pre-training with over 500 million tokens of legal texts. Our innovative approach significantly improves capabilities in legal tasks by using Large Language Models (LLMs) to convert raw training data into reading comprehension text. Our legal LLMs have demonstrated superior performance in legal benchmarks, even outperforming models trained on much larger datasets with more resources. This work emphasizes the effectiveness of continued pre-training on domain-specific texts, while using affordable LLMs for data conversion, which gives these models domain expertise while retaining general language understanding capabilities. While this work uses the legal domain as a test case, our method can be scaled and applied to any pre-training dataset, resulting in significant improvements across different tasks. These findings underscore the potential of domain-adaptive pre-training and reading comprehension for the development of highly effective domain-specific language models.

法律AI持续预训练数据生成小模型大用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。