arXiv:2504.07087cs.CLcs.AI2025-04被引 8

构建可扩展基准,评估大模型在文本化知识图谱上的推理能力

KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs

  • 设计五类知识图谱理解任务的评测基准
  • 对比七种模型与五种文本化策略的效果差异
  • 为提升大模型知识推理性能提供实证指导

知识图谱已成为向大语言模型注入最新事实知识的常用方法,通常通过将知识图谱转化为大模型可处理的文本形式实现。尽管已有多种知识图谱编码方法被提出,但这种文本化过程对大模型性能的影响仍缺乏深入研究。本文提出KG-LLM-Bench,一个涵盖五类知识图谱理解任务的全面且可扩展的评测基准,并评估不同编码策略在多种基础模型上的表现。通过对七种语言模型和五种文本化策略的广泛实验,揭示了优化大模型在知识图谱推理任务中性能的关键因素。

原文摘要 · Abstract (English)

Knowledge graphs have emerged as a popular method for injecting up-to-date, factual knowledge into large language models (LLMs). This is typically achieved by converting the knowledge graph into text that the LLM can process in context. While multiple methods of encoding knowledge graphs have been proposed, the impact of this textualization process on LLM performance remains under-explored. We introduce KG-LLM-Bench, a comprehensive and extensible benchmark spanning five knowledge graph understanding tasks, and evaluate how different encoding strategies affect performance across various base models. Our extensive experiments with seven language models and five textualization strategies provide insights for optimizing LLM performance on KG reasoning tasks.

知识图谱大模型推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。