arXiv:2502.05239cs.CLcs.AI2025-02被引 2

改进知识图谱构建评估,重点减少幻觉和遗漏问题。

Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics

  • 引入BERTScore提升图相似度评估,设95%为匹配阈值。
  • 微调后的Mistral模型显著降低幻觉与遗漏,准确率提升。
  • 适合关注知识图谱质量评估与模型可靠性的研究者。

大语言模型在从非结构化文本自动构建知识图谱方面展现出巨大潜力。本文基于先前工作[16],在精确率、召回率、F1分数、三元组匹配和图匹配等指标基础上,提出一种改进的评估框架,重点关注幻觉与遗漏问题。我们引入BERTScore作为图相似度度量,并设定95%的图匹配实用阈值。实验聚焦Mistral模型,对比其原始版与微调版在零样本和少样本设置下的表现。进一步使用KELM-sub训练数据集中的样例进行扩展实验,结果表明微调模型显著提升了知识图谱构建的准确性,同时降低了精确幻觉与遗漏率。但研究也发现,微调模型在KELM-sub数据集上的泛化能力反而下降。本研究强调了综合评估指标对推动文本到知识图谱构建技术进步的重要性。

原文摘要 · Abstract (English)

Recent advancements in large language models have demonstrated significant potential in the automated construction of knowledge graphs from unstructured text. This paper builds upon our previous work [16], which evaluated various models using metrics like precision, recall, F1 score, triple matching, and graph matching, and introduces a refined approach to address the critical issues of hallucination and omission. We propose an enhanced evaluation framework incorporating BERTScore for graph similarity, setting a practical threshold of 95% for graph matching. Our experiments focus on the Mistral model, comparing its original and fine-tuned versions in zero-shot and few-shot settings. We further extend our experiments using examples from the KELM-sub training dataset, illustrating that the fine-tuned model significantly improves knowledge graph construction accuracy while reducing the exact hallucination and omission. However, our findings also reveal that the fine-tuned models perform worse in generalization tasks on the KELM-sub dataset. This study underscores the importance of comprehensive evaluation metrics in advancing the state-of-the-art in knowledge graph construction from textual data.

知识图谱大模型评估幻觉检测图匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。