arXiv:2505.24478cs.AIcs.CL2025-05被引 16

调优知识图谱与大模型接口,显著提升复杂推理准确率

Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning

  • 针对分块、图构建、检索等环节系统调参
  • 在多跳问答任务中提升精确匹配与正确性指标
  • 适合需要高精度推理的系统开发者参考

将大语言模型(LLMs)与知识图谱(KGs)结合会形成包含大量超参数的复杂系统,直接影响性能。尽管这类系统在检索增强生成中日益普遍,但系统的超参数优化仍缺乏系统研究。本文以Cognee这一端到端知识图谱构建与检索的模块化框架为背景,基于三个多跳问答基准(HotPotQA、TwoWikiMultiHop、MuSiQue),对分块、图构建、检索和提示工程等参数进行优化。每个配置使用精确匹配、F1值以及DeepEval的基于LLM的正确性评估指标进行评分。结果表明,针对性调优可带来显著性能提升,但增益不一致,不同数据集和指标间表现差异明显。该现象既凸显了调优的价值,也揭示了现有评估方法的局限性。本文强调,未来进展不仅依赖架构创新,更需建立清晰的优化与评估框架,以应对复杂模块化系统的需求。

原文摘要 · Abstract (English)

Integrating Large Language Models (LLMs) with Knowledge Graphs (KGs) results in complex systems with numerous hyperparameters that directly affect performance. While such systems are increasingly common in retrieval-augmented generation, the role of systematic hyperparameter optimization remains underexplored. In this paper, we study this problem in the context of Cognee, a modular framework for end-to-end KG construction and retrieval. Using three multi-hop QA benchmarks (HotPotQA, TwoWikiMultiHop, and MuSiQue) we optimize parameters related to chunking, graph construction, retrieval, and prompting. Each configuration is scored using established metrics (exact match, F1, and DeepEval's LLM-based correctness metric). Our results demonstrate that meaningful gains can be achieved through targeted tuning. While the gains are consistent, they are not uniform, with performance varying across datasets and metrics. This variability highlights both the value of tuning and the limitations of standard evaluation measures. While demonstrating the immediate potential of hyperparameter tuning, we argue that future progress will depend not only on architectural advances but also on clearer frameworks for optimization and evaluation in complex, modular systems.

知识图谱大模型推理优化多跳问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。