用大模型解决法律文本中人物指代模糊问题,构建更清晰的走私网络图谱
LINK-KG: LLM-Driven Coreference-Resolved Knowledge Graphs for Human Smuggling Networks
- 分三阶段用大模型识别并统一法律文档中的指代关系
- 节点重复率降45.21%,噪声节点减32.22%,图谱更干净
- 适合司法分析、犯罪网络研究等需要精准实体链接的场景
人类走私网络结构复杂且持续演变,难以全面分析。法律案件文档虽包含丰富事实与程序信息,但普遍篇幅长、无结构,且存在模糊或动态变化的指代问题,给自动化知识图谱(KG)构建带来挑战。现有方法或忽略指代消解,或无法扩展至长文本,导致图谱碎片化、实体链接不一致。我们提出LINK-KG,一个模块化框架,将三阶段大模型引导的指代消解流程与下游知识图谱抽取结合。核心是类型特定的提示缓存(type-specific Prompt Cache),可在文档分块间持续追踪并解析指代,实现从短文到长文的清晰、去歧化叙事,支持结构化知识图谱构建。相比基线方法,LINK-KG平均节点重复率降低45.21%,噪声节点减少32.22%,显著提升图谱清洁度与连贯性,为复杂犯罪网络分析提供坚实基础。
原文摘要 · Abstract (English)
Human smuggling networks are complex and constantly evolving, making them difficult to analyze comprehensively. Legal case documents offer rich factual and procedural insights into these networks but are often long, unstructured, and filled with ambiguous or shifting references, posing significant challenges for automated knowledge graph (KG) construction. Existing methods either overlook coreference resolution or fail to scale beyond short text spans, leading to fragmented graphs and inconsistent entity linking. We propose LINK-KG, a modular framework that integrates a three-stage, LLM-guided coreference resolution pipeline with downstream KG extraction. At the core of our approach is a type-specific Prompt Cache, which consistently tracks and resolves references across document chunks, enabling clean and disambiguated narratives for structured knowledge graph construction from both short and long legal texts. LINK-KG reduces average node duplication by 45.21% and noisy nodes by 32.22% compared to baseline methods, resulting in cleaner and more coherent graph structures. These improvements establish LINK-KG as a strong foundation for analyzing complex criminal networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。