提升法律文本知识图谱构建精度,减少重复节点与噪声。
Inside CORE-KG: Evaluating Structured Prompting and Coreference Resolution for Knowledge Graphs
- 引入类型感知共指消解与领域引导的结构化提示
- 移除共指消解使节点重复率升28.25%,噪声增4.32%
- 移除结构化提示导致噪声激增73.33%,适合法律信息抽取场景
人贩子网络日益复杂,法律案件文档虽具关键价值,但常为非结构化、词汇密集且存在模糊或变化的指代,给自动化知识图谱(KG)构建带来挑战。现有基于大模型的方法虽优于静态模板,但仍生成噪声多、碎片化的图谱,因缺乏引导提取与共指消解。新提出的CORE-KG框架通过整合类型感知共指模块和领域引导的结构化提示,显著降低节点重复与法律噪声。本文对CORE-KG进行系统消融研究,量化其两大核心组件的贡献。结果表明:移除共指消解导致节点重复率上升28.25%、噪声节点增加4.32%;移除结构化提示则使节点重复率上升4.29%、噪声节点飙升73.33%。这些实证发现为从复杂法律文本中构建鲁棒大模型管线提供了重要依据。
原文摘要 · Abstract (English)
Human smuggling networks are increasingly adaptive and difficult to analyze. Legal case documents offer critical insights but are often unstructured, lexically dense, and filled with ambiguous or shifting references, which pose significant challenges for automated knowledge graph (KG) construction. While recent LLM-based approaches improve over static templates, they still generate noisy, fragmented graphs with duplicate nodes due to the absence of guided extraction and coreference resolution. The recently proposed CORE-KG framework addresses these limitations by integrating a type-aware coreference module and domain-guided structured prompts, significantly reducing node duplication and legal noise. In this work, we present a systematic ablation study of CORE-KG to quantify the individual contributions of its two key components. Our results show that removing coreference resolution results in a 28.25% increase in node duplication and a 4.32% increase in noisy nodes, while removing structured prompts leads to a 4.29% increase in node duplication and a 73.33% increase in noisy nodes. These findings offer empirical insights for designing robust LLM-based pipelines for extracting structured representations from complex legal texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。