用大模型构建可浏览审计的消歧知识库,支持自然语言查询
GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base

- 通过上下文引导消歧,递归构建去重合并的实体关系网络
- 包含3840万三元组、160万实体、207600种关系和66000类
- 支持事实溯源、自然语言转查询、文本链接到标准实体
我们展示了一个基于大语言模型生成的大规模消歧知识库(GPTKB 2.0)的网页演示。该知识库包含3840万条三元组,覆盖160万条规范实体,以及207600种合并后的关系和66000种合并后的类别。与以往依赖表面字符串识别实体的LLM衍生知识库不同,GPTKB 2.0在递归构建过程中进行上下文引导的消歧,将同音异义词分离、同义提及合并。演示界面支持用户浏览实体、跨知识库跳转,审计单个事实的来源,包括原始表达式、候选匹配、源三元组及消歧决策。还支持结构化SPARQL查询、自然语言问题转SPARQL,以及从用户输入文本中实体链接至标准的GPTKB 2.0条目。GPTKB 2.0可在https://gptkb.org/访问,完整知识库可下载离线使用。
原文摘要 · Abstract (English)
We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities, together with 207.6K consolidated relations and 66K consolidated classes. Unlike prior LLM-derived knowledge bases that largely identify entities by surface strings, GPTKB 2.0 performs context-guided disambiguation during recursive KB construction, separating homonyms and merging synonymous mentions as facts are elicited. The demo makes this process inspectable: users can browse entities, follow links across the KB, and audit the provenance of individual facts, including surface forms, candidate matches, source triples, and disambiguation decisions. The interface further supports structured SPARQL queries, natural-language questions translated to SPARQL, and entity linking from user-provided text to canonical GPTKB 2.0 entries. GPTKB 2.0 is available at https://gptkb.org/, with the full KB downloadable for offline use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。