用智能代理提升历史文献检索准确率,抗错防幻觉。
Robust Interpretation of Historical Documents in Knowledge Graphs Through Query Inference and Execution

- 设计智能体系统,融合关键词定位与知识图谱生成查询
- 在存在识别错误时仍保持高准确率,优于传统RAG方法
- 适合数字档案馆、历史研究者使用,提升可验证性
大型语言模型(LLMs)重塑了用户与数字信息的交互方式,但其广泛且随意的集成引发了可靠性与可信度问题,尤其在数字图书馆和历史档案中更为突出。如何在利用LLM泛化能力的同时,维持档案机构所需的责任感?本文提出一种智能体检索系统,旨在实现更准确、可验证的历史数据访问,同时保留无约束LLM的灵活性。作为历史文献分析的贡献,我们比较了传统RAG与基于知识图谱的Agentic GraphRAG在真实场景下的表现,包括光学字符识别(OCR)和转录错误的存在。提出一种半符号框架,结合词位定位技术进行后处理纠错,并通过知识图谱支持合成查询的生成。词位定位与代码生成的交替协作,使智能体能构建对误读和幻觉鲁棒的检索查询,同时在噪声和不确定性普遍存在的历史文献分析中,仍可借助近似搜索实现精准定位。
原文摘要 · Abstract (English)
The emergence of Large Language Models (LLMs) has redefined how users interact with information in digital environments. However, their widespread and often indiscriminate integration has raised significant concerns regarding reliability and trustworthiness issues that are particularly critical when accessing digital libraries and historical archives. How can one leverage the generalization capacity of an LLM without losing the level of accountability required for an archival institution? In this paper, we present an agentic retrieval system designed to deliver more accurate and verifiable access to historical data while preserving much of the flexibility associated with unconstrained LLMs. As a contribution to historical document analysis, we compare traditional Retrieval-Augmented Generation (RAG) with an agentic GraphRAG architecture in their ability to deliver historical information under realistic conditions, including the presence of OCR and transcription errors. We introduce a semi-symbolic framework that integrates word-spotting techniques for post-OCR correction with a knowledge graph representation that enables the agent to access information through synthesized queries. The interleaved collaboration between word spotting and code generation allows the agent to construct strong retrieval queries that are robust to misinterpretation and hallucination, while still leveraging approximate search when noise and uncertainty, common in historical document analysis, would otherwise hinder precise retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。