用GPT-4.1构建1亿三元组知识库,可探索模型事实知识的盲区。
GPTKB v1.5: A Massive Knowledge Base for Exploring Factual LLM Knowledge
- 通过递归调用GPT-4.1自动抽取并链接事实,构建1亿三元组知识库
- 耗资1.4万美元,实现对大模型知识的结构化存储与可查性
- 适合研究者分析模型知识边界,或用于知识图谱自动化构建
语言模型虽强大,但其事实知识仍难以理解且无法随意浏览或进行大规模统计分析。本演示介绍GPTKB v1.5——一个基于GPT-4.1、花费1.4万美元构建的密集互联1亿三元组知识库(KB),采用GPTKB方法实现大规模递归式大模型知识实体化。演示聚焦三个应用场景:(1) 基于链路遍历的LLM知识探索,(2) 基于SPARQL的结构化知识查询,(3) 对比分析大模型知识的优势与局限。大规模递归式大模型知识实体化为系统性分析模型知识及自动化知识库构建提供了突破性机遇。
原文摘要 · Abstract (English)
Language models are powerful artifacts, yet their factual knowledge is still poorly understood, and inaccessible to ad-hoc browsing and scalable statistical analysis. This demonstration introduces GPTKB v1.5, a densely interlinked 100-million-triple knowledge base (KB) built for $14,000 from GPT-4.1, using the GPTKB methodology for massive-recursive LLM knowledge materialization. This demo focuses on three use cases: (1) link-traversal-based LLM knowledge exploration, (2) SPARQL-based structured LLM knowledge querying, (3) comparative exploration of the strengths and weaknesses of LLM knowledge. Massive-recursive LLM knowledge materialization is a groundbreaking opportunity both for the systematic analysis of LLM knowledge, as well as for automated KB construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。