arXiv:2504.16414cs.CL2025-04被引 3

用化学领域知识图谱测试大模型多跳推理能力,发现现有模型仍存显著短板。

Evaluating Multi-Hop Reasoning in Large Language Models: A Chemistry-Centric Case Study

  • 构建化学知识图谱,自动生成多跳推理问题
  • 顶尖模型在无外部信息时准确率不足50%
  • 适合研究大模型推理机制与知识增强的学者

本研究提出一个聚焦化学领域的全新基准,包含精心筛选的数据集和评估流程,用于检验大语言模型的组合式推理能力。我们设计并验证了一套全自动管道,经领域专家确认,通过结合OpenAI推理模型与命名实体识别(NER)系统,从近期文献中提取化学实体,并借助外部知识库构建全面的知识图谱。基于该图谱生成多跳问题,在有无上下文增强两种条件下评估模型表现。实验表明,即使最先进的模型在多跳推理任务中仍面临重大挑战。结果强调了引入文档检索对提升性能的关键作用,即便实现完美检索,仍无法完全消除推理错误,凸显组合推理的复杂性。该工作不仅揭示了当前大模型的局限,还提供了一种可推广至其他领域的挑战性数据生成新方法,推动计算语言学中推理能力的理解。

原文摘要 · Abstract (English)

In this study, we introduced a new benchmark consisting of a curated dataset and a defined evaluation process to assess the compositional reasoning capabilities of large language models within the chemistry domain. We designed and validated a fully automated pipeline, verified by subject matter experts, to facilitate this task. Our approach integrates OpenAI reasoning models with named entity recognition (NER) systems to extract chemical entities from recent literature, which are then augmented with external knowledge bases to form a comprehensive knowledge graph. By generating multi-hop questions across these graphs, we assess LLM performance in both context-augmented and non-context augmented settings. Our experiments reveal that even state-of-the-art models face significant challenges in multi-hop compositional reasoning. The results reflect the importance of augmenting LLMs with document retrieval, which can have a substantial impact on improving their performance. However, even perfect retrieval accuracy with full context does not eliminate reasoning errors, underscoring the complexity of compositional reasoning. This work not only benchmarks and highlights the limitations of current LLMs but also presents a novel data generation pipeline capable of producing challenging reasoning datasets across various domains. Overall, this research advances our understanding of reasoning in computational linguistics.

多跳推理知识图谱化学信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。