arXiv:2604.14220cs.IRcs.AI2026-04

用智能爬取构建企业文档知识图谱,让复杂法规查询更准更快

Knowledge Graph RAG: Agentic Crawling and Graph Construction in Enterprise Documents

论文配图:Knowledge Graph RAG: Agentic Crawling and Graph Construction in Enterprise Documents
图 1 · 摘自论文原文
  • 设计代理式爬取机制,自动挖掘文档间多跳引用关系
  • 在联邦法规数据集上比传统检索系统准确率提升70%
  • 适合需要精准理解复杂规章的企业用户

本文针对企业文档体系中语义搜索的局限性提出解决方案。传统RAG管道难以捕捉层级化与关联性信息,导致检索不准确。我们提出基于代理的知識圖譜(Agentic Knowledge Graphs),采用递归爬取机制有效应对嵌套逻辑与多跳引用。在《联邦法规》(Code of Federal Regulations, CFR)上的基准测试表明,该知识图谱增强方法相比标准向量检索RAG系统,准确率提升70%,能为复杂监管查询提供全面且精确的答案。

原文摘要 · Abstract (English)

This research paper addresses the limitations of semantic search in complex enterprise document ecosystems. Traditional RAG pipelines often fail to capture hierarchical and interconnected information, leading to retrieval inaccuracies. We propose Agentic Knowledge Graphs featuring Recursive Crawling as a robust solution for navigating superseding logic and multi-hop references. Our benchmark evaluation using the Code of Federal Regulations (CFR) demonstrates that this Knowledge Graph-enhanced approach achieves a 70% accuracy improvement over standard vector-based RAG systems, providing exhaustive and precise answers for complex regulatory queries.

知识图谱RAG企业文档智能爬取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。