arXiv:2504.12422cs.HCcs.AI2025-04被引 4

用知识图谱约束大模型,减少幻觉,提升问答可信度。

Mitigating LLM Hallucinations with Knowledge Graphs: A Case Study

  • 让大模型问答前必须查询知识图谱获取真实数据
  • 在标准测试集上表现优于GPT-4,但部分问题仍答不准
  • 适合安全、医疗等对准确性要求高的场景

高风险领域如网络作战需要可靠可信的AI方法。尽管大语言模型(LLMs)在这些领域日益普及,但仍存在幻觉问题。本文基于一个名为LinkQ的开源自然语言接口开展案例研究,该系统通过强制大模型在问答时查询知识图谱(KG)获取真实数据来缓解幻觉。我们使用一个知名的知识图谱问答(KGQA)数据集对LinkQ进行量化评估,结果显示其性能优于GPT-4,但在某些问题类别上仍表现不佳,表明未来需探索更优的查询构建策略。此外,我们还通过两名领域专家对真实网络安全知识图谱进行定性研究,总结出专家反馈、改进建议、系统局限及未来发展方向。

原文摘要 · Abstract (English)

High-stakes domains like cyber operations need responsible and trustworthy AI methods. While large language models (LLMs) are becoming increasingly popular in these domains, they still suffer from hallucinations. This research paper provides learning outcomes from a case study with LinkQ, an open-source natural language interface that was developed to combat hallucinations by forcing an LLM to query a knowledge graph (KG) for ground-truth data during question-answering (QA). We conduct a quantitative evaluation of LinkQ using a well-known KGQA dataset, showing that the system outperforms GPT-4 but still struggles with certain question categories - suggesting that alternative query construction strategies will need to be investigated in future LLM querying systems. We discuss a qualitative study of LinkQ with two domain experts using a real-world cybersecurity KG, outlining these experts' feedback, suggestions, perceived limitations, and future opportunities for systems like LinkQ.

知识图谱大模型幻觉抑制可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。