arXiv:2506.21607cs.CLcs.AI2025-06被引 2

用大模型构建更清晰的人口走私网络知识图谱

CORE-KG: An LLM-Driven Knowledge Graph Construction Framework for Human Smuggling Networks

  • 分两步走:先解决指代歧义,再精准提取实体关系
  • 节点重复减少33.28%,法律文本噪声降低38.37%
  • 适合司法分析、犯罪网络研究者使用

人口走私网络日益复杂且具适应性,法律案件文档虽含关键信息,但内容非结构化、词汇密集且指代模糊,给自动化知识图谱(KG)构建带来挑战。现有方法多依赖静态模板,缺乏指代消解能力;近期基于大模型的方法常因幻觉产生噪声和碎片化图谱,且因缺乏引导导致节点重复。我们提出CORE-KG,一个模块化框架,用于从法律文本中构建可解释的知识图谱。该框架采用两阶段流程:(1) 通过序列化、结构化的提示,实现类型感知的指代消解;(2) 基于领域引导指令,在改进的GraphRAG框架上进行实体与关系抽取。相比基于GraphRAG的基线,CORE-KG将节点重复率降低33.28%,法律文本噪声减少38.37%,生成的图谱更清晰、连贯。该成果为复杂犯罪网络分析提供了可靠基础。

原文摘要 · Abstract (English)

Human smuggling networks are increasingly adaptive and difficult to analyze. Legal case documents offer valuable insights but are unstructured, lexically dense, and filled with ambiguous or shifting references-posing challenges for automated knowledge graph (KG) construction. Existing KG methods often rely on static templates and lack coreference resolution, while recent LLM-based approaches frequently produce noisy, fragmented graphs due to hallucinations, and duplicate nodes caused by a lack of guided extraction. We propose CORE-KG, a modular framework for building interpretable KGs from legal texts. It uses a two-step pipeline: (1) type-aware coreference resolution via sequential, structured LLM prompts, and (2) entity and relationship extraction using domain-guided instructions, built on an adapted GraphRAG framework. CORE-KG reduces node duplication by 33.28%, and legal noise by 38.37% compared to a GraphRAG-based baseline-resulting in cleaner and more coherent graph structures. These improvements make CORE-KG a strong foundation for analyzing complex criminal networks.

知识图谱大模型犯罪分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。