arXiv:2605.18770cs.IRcs.AI2026-05被引 1

用图结构提升商业注册数据的可审计分析,准确率从26%升至83%

Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis

论文配图:Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis
图 1 · 摘自论文原文
  • 构建知识图谱融合结构化字段与非结构化文本,实现实体精准关联
  • 在瑞士商业公报数据上,事实正确率从0.26提升至0.83,相关性与完整性同步改善
  • 支持可追溯的对话分析,适合审计、合规等需要透明推理的场景

公开商业注册信息虽形式开放,但实际分析困难,因关键事实分散于数百万条记录中,涵盖结构化元数据、多语言法律通知、时间事件及实体别名。本文提出一种受控的工具驱动型图RAG架构,用于可审计的自然语言分析。该系统将瑞士官方商业公报内容转换为包含超五百万节点和470万关系的Neo4j知识图谱,整合结构化字段的确定性注入、基于大模型的未明示主体抽取,以及确定性的身份解析层。分析代理通过意图路由、受限图工具、有界反思和状态机引导的回答生成,在图谱上执行操作。采用多层级评估协议,涵盖答案质量、检索行为、实体解析和多轮对话表现。与密集、词法、混合的扁平检索基线及架构消融实验对比,在人工标注基准上,图媒介检索使事实正确率由最强基线的0.26提升至0.83,相关性与完整性亦显著改善。消融实验表明,有界反思提升答案质量,意图路由与大模型图增强则提高复杂实体解析任务的可靠性。探索性仪表板展示每项回答背后的图证据与执行轨迹,支持用户审查推理过程。

原文摘要 · Abstract (English)

Public commercial registries are formally open, yet their practical analysis remains difficult because relevant facts are scattered across millions of records that combine structured metadata, multilingual legal notices, temporal events, and entity aliases. This paper presents a controlled, tool-mediated agentic GraphRAG architecture for auditable natural-language analysis of such registries. The proposed pipeline transforms publications from the Swiss Official Gazette of Commerce into a Neo4j knowledge graph comprising over five million nodes and 4.7 million relationships. It combines deterministic ingestion of structured registry fields, LLM-assisted extraction of latent actors from unstructured notices, and a deterministic identity-resolution layer. An analytical agent operates on this graph through intent routing, restricted graph tools, bounded reflection, and state-machine-guided response synthesis. We evaluate the system using a multi-tier protocol covering answer quality, retrieval behavior, entity resolution, and multi-turn conversational performance. The complete architecture is compared with dense, lexical, and hybrid flat-retrieval baselines and with controlled architectural ablations. On a manually curated benchmark, graph-mediated retrieval increases factual correctness from 0.26 for the strongest flat-retrieval baseline to 0.83 for the complete system, with comparable improvements in relevance and completeness. Ablation results show that bounded reflection improves answer quality while intent routing and LLM-based graph enrichment improve reliability in difficult entity resolution tasks. An exploratory dashboard displays the graph evidence and execution traces underlying each response, allowing users to inspect how answers were produced.

图神经网络可解释性知识图谱实体对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。