arXiv:2602.03417cs.CL2026-02

构建百亿规模多语言事实图谱,解决大模型幻觉与跨语言知识对齐问题。

FactNet: A Billion-Scale Knowledge Graph for Multilingual Factual Grounding

  • 将17亿个维基数据断言与301亿条多语种维基来源证据绑定
  • 在跨语言知识补全任务中,多语言结构提升不同语言间知识迁移效果
  • 提供含防泄漏机制的评测基准,支持多任务评估

大型语言模型存在编造事实且难以追溯依据的问题,尤其在非英语语言中更为突出。现有资源存在权衡:结构化知识库缺乏文本证据,而有证据的数据集规模小且仅限单语。我们提出FactNet,一个百亿级开放资源,将17亿个维基数据断言与来自316个本土维基百科版本的301亿条证据指针相耦合。FactNet采用确定性构建流程,确保每条证据可追溯至原始文本的字节级位置。我们进一步建立FactNet-Bench评测基准,涵盖知识图谱补全、问答与事实核查任务,并配备系统性泄漏控制机制。实验表明,FactNet-Bench能有效区分结构型、文本感知型及大模型融合型方法,且跨语言结构显著促进不同语言层级间的知识迁移。

原文摘要 · Abstract (English)

Large language models hallucinate factual claims and struggle to ground their outputs in retrievable evidence, particularly in non-English languages. Existing resources impose a trade-off: structured knowledge bases lack textual grounding, whereas grounded datasets remain small and monolingual. We introduce FactNet, a billion-scale open resource that couples 1.7B Wikidata assertions with 3.01B evidence pointers drawn from 316 native Wikipedia editions. FactNet employs a deterministic construction pipeline, ensuring that every evidence unit is traceable to its source with byte-level precision. We further establish FactNet-Bench, an evaluation suite for Knowledge Graph Completion, Question Answering, and Fact Checking, equipped with systematic leakage controls. Experiments demonstrate that FactNet-Bench differentiates among structural, text-aware, and LLM-integrated methods, and that cross-lingual structure enables knowledge transfer across language tiers.

知识图谱多语言事实核查大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。