arXiv:2505.14101cs.CL2025-05被引 5

构建多语言知识图谱基准,评估大模型幻觉问题。

MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations

  • 基于知识图谱构建多语言多跳推理数据集。
  • 融合知识图谱使模型幻觉检测提升0.29至0.42分。
  • 适合研究大模型事实性与跨语言幻觉检测者使用。

大语言模型存在忠实性与事实性缺陷,常表现为幻觉。现有评测基准多聚焦英文语境,依赖网页链接或文本片段,忽略结构化事实资源。知识图谱(KG)因其结构化表示实体关系的特点,可有效缓解幻觉。为此,我们提出多语言、多跳的知识图谱驱动评测基准 MultiHal,用于生成式文本的事实性评估。通过从开放域知识图谱中挖掘14万条路径,并经去噪处理,保留高质量的25.9万条路径。基线评估显示,相较于纯问答任务,采用知识图谱增强检索(KG-RAG)后,语义相似度提升0.12至0.36,自然语言推理蕴含度提升0.16至0.36,幻觉检测准确率提升0.29至0.42,验证了知识图谱集成的有效性。我们期望 MultiHal 能推动面向图结构的幻觉缓解与事实核查研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have inherent limitations of faithfulness and factuality, commonly referred to as hallucinations. Several benchmarks have been developed that provide a test bed for factuality evaluation within the context of English-centric datasets, while relying on supplementary informative context like web links or text passages but ignoring the available structured factual resources. To this end, Knowledge Graphs (KGs) have been identified as a useful aid for hallucination mitigation, as they provide a structured way to represent the facts about entities and their relations with minimal linguistic overhead. We bridge the lack of KG paths and multilinguality for factual language modeling within the existing hallucination evaluation benchmarks and propose a KG-based multilingual, multihop benchmark called MultiHal framed for generative text evaluation. As part of our data collection pipeline, we mined 140k KG-paths from open-domain KGs, from which we pruned noisy KG-paths, curating a high-quality subset of 25.9k. Our baseline evaluation shows an absolute scale improvement by approximately 0.12 to 0.36 points for the semantic similarity score, 0.16 to 0.36 for NLI entailment and 0.29 to 0.42 for hallucination detection in KG-RAG over vanilla QA across multiple languages and multiple models, demonstrating the potential of KG integration. We anticipate MultiHal will foster future research towards several graph-based hallucination mitigation and fact-checking tasks.

大模型幻觉检测知识图谱多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。