用领域微调模型提升走私网络知识图谱构建效率与质量
FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs
- 基于微调大模型实现法律文本中的实体与关系抽取
- 实体和关系准确率分别提升15.50%和31.46%
- 处理速度加快50%,噪声减少近半,适合司法情报分析
法庭文件包含关于人口走私网络的重要证据,但这些信息常隐藏在非结构化、术语密集的法律文档中。尽管大语言模型(LLMs)可通过自动化信息抽取支持知识图谱构建,现有方法多依赖通用模型,无法适配该领域的实体与关系定义。我们提出FineREX,一个围绕微调大模型构建的命名实体识别与关系抽取(NER-RE)的轻量级知识图谱构建流程。基于512个文本片段的手动标注数据集,FineREX相比更大规模通用基线,在实体和关系的F1得分上分别提升15.50%和31.46%。结果生成的知识图谱显著降低法律噪声近一半,长文档中的节点重复率从17.78%降至11.17%。通过省去文档重写和冗余抽取环节,端到端处理时间减少50.0%。实验表明,领域微调可显著超越更大通用模型,同时提升非法网络分析中知识图谱的质量与效率。
原文摘要 · Abstract (English)
Court proceedings contain valuable evidence about human smuggling networks, but this information is often buried within unstructured, jargon-heavy legal documents. While large language models (LLMs) can support knowledge graph construction through automated information extraction, existing approaches rely on general-purpose models that are not tailored to the entity and relationship definitions required in this domain. We introduce FineREX, a streamlined knowledge graph construction pipeline built around a fine-tuned LLM for named entity recognition and relationship extraction (NER-RE). Using a manually annotated dataset of $512$ text chunks, FineREX achieves absolute improvements of 15.50% and 31.46% in entity and relationship F1-score, respectively, compared to a larger general-purpose baseline. These gains translate into higher-quality knowledge graphs, reducing legal noise by nearly half and lowering node duplication on long documents from 17.78% to 11.17%. By eliminating document rewriting and redundant extraction stages, FineREX also reduces end-to-end processing time by 50.0%. Our results demonstrate that domain-specific fine-tuning can substantially outperform larger general-purpose models while improving both the quality and efficiency of knowledge graph construction for illicit network analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。