用大模型构建高质量合同数据集,提升工业合同信息提取效果
Large Language Model for Extracting Complex Contract Information in Industrial Scenes
- 基于聚类与GPT-4/3.5生成高质量标注数据
- 通过关键词组合增强数据,提升模型鲁棒性
- 结合LoRA与数据平衡,实现高召回率与精度
本文提出一种面向工业场景复杂合同信息抽取任务的高质量数据集构建方法,并在此基础上微调大型语言模型。首先对工业合同文本进行聚类分析,利用GPT-4和GPT-3.5从原始合同数据中提取关键信息,获得高质量数据标注;其次通过构建新文本实现数据增强,使用GPT-3.5从随机组合的关键词生成非结构化合同文本,提升模型鲁棒性;最后基于高质量数据集微调大模型。实验结果表明,该模型在保证高领域召回率与精确率的同时,具备优异的整体性能,并兼顾解析效率。LoRA、数据平衡与数据增强显著提升了模型准确率与鲁棒性。所提方法为工业合同信息抽取任务提供了新颖高效的解决方案。
原文摘要 · Abstract (English)
This paper proposes a high-quality dataset construction method for complex contract information extraction tasks in industrial scenarios and fine-tunes a large language model based on this dataset. Firstly, cluster analysis is performed on industrial contract texts, and GPT-4 and GPT-3.5 are used to extract key information from the original contract data, obtaining high-quality data annotations. Secondly, data augmentation is achieved by constructing new texts, and GPT-3.5 generates unstructured contract texts from randomly combined keywords, improving model robustness. Finally, the large language model is fine-tuned based on the high-quality dataset. Experimental results show that the model achieves excellent overall performance while ensuring high field recall and precision and considering parsing efficiency. LoRA, data balancing, and data augmentation effectively enhance model accuracy and robustness. The proposed method provides a novel and efficient solution for industrial contract information extraction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。