用大模型检测维基百科中的事实矛盾,提升知识准确性。
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
- 构建智能系统CLAIRE,结合大模型推理与检索发现潜在矛盾。
- 实测显示编辑使用后识别矛盾效率提高64.7%,信心提升87.5%。
- 发现3.3%的维基条目存在矛盾,适合知识审核与模型评估者参考。
维基百科是全球最大的开放知识库,广泛用于训练大语言模型(LLMs)和检索增强生成(RAG)系统,其准确性至关重要。本文聚焦事实不一致这一特定错误类型,提出跨语料库不一致检测任务。我们设计了CLAIRE——一个结合大模型推理与检索的智能系统,可发现潜在矛盾并提供上下文证据供人工审核。在与资深维基编辑的用户研究中,87.5%的参与者表示使用CLAIRE后信心更高,且在同一时间内识别出的不一致项增加了64.7%。结合人工标注,我们构建了首个真实维基不一致基准数据集WIKICOLLIDE。通过随机采样与CLAIRE辅助分析,发现至少3.3%的英文维基事实与其他事实矛盾,且这些矛盾已影响到7.3%的FEVEROUS和4.0%的AmbigQA样本。在该数据集上基准测试表明,当前最佳全自动系统仅达到75.1% AUROC,仍有巨大改进空间。结果表明,矛盾是维基知识中可度量的部分,而类似CLAIRE的大模型系统可为编辑提供规模化提升知识一致性的实用工具。
原文摘要 · Abstract (English)
Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RAG) systems. Ensuring its accuracy is therefore critical. But how accurate is Wikipedia, and how can we improve it? We focus on inconsistencies, a specific type of factual inaccuracy, and introduce the task of corpus-level inconsistency detection. We present CLAIRE, an agentic system that combines LLM reasoning with retrieval to surface potentially inconsistent claims along with contextual evidence for human review. In a user study with experienced Wikipedia editors, 87.5% reported higher confidence when using CLAIRE, and participants identified 64.7% more inconsistencies in the same amount of time. Combining CLAIRE with human annotation, we contribute WIKICOLLIDE, the first benchmark of real Wikipedia inconsistencies. Using random sampling with CLAIRE-assisted analysis, we find that at least 3.3% of English Wikipedia facts contradict another fact, with inconsistencies propagating into 7.3% of FEVEROUS and 4.0% of AmbigQA examples. Benchmarking strong baselines on this dataset reveals substantial headroom: the best fully automated system achieves an AUROC of only 75.1%. Our results show that contradictions are a measurable component of Wikipedia and that LLM-based systems like CLAIRE can provide a practical tool to help editors improve knowledge consistency at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。