arXiv:2412.08317cs.CL2024-12被引 2

大模型在多跳推理中仍难有效利用外部知识

Large Language Models Still Face Challenges in Multi-Hop Reasoning with External Knowledge

  • 测试GPT-3.5在四种链式思考提示下的多跳推理能力
  • 模型在多跳任务中表现远低于人类,尤其在复杂推理时
  • 适合研究大模型推理缺陷与提升方法的学者

我们通过一系列实验,从三个方面评估大语言模型的多跳推理能力:选择并组合外部知识、处理非顺序推理任务,以及推广到更多跳数的数据样本。我们在四个推理基准上测试了GPT-3.5模型,采用链式思考提示(Chain-of-Thought prompting)及其变体。结果表明,尽管大语言模型在多种推理任务中表现优异,但在多跳推理方面仍存在严重缺陷,与人类表现差距显著。

原文摘要 · Abstract (English)

We carry out a series of experiments to test large language models' multi-hop reasoning ability from three aspects: selecting and combining external knowledge, dealing with non-sequential reasoning tasks and generalising to data samples with larger numbers of hops. We test the GPT-3.5 model on four reasoning benchmarks with Chain-of-Thought prompting (and its variations). Our results reveal that despite the amazing performance achieved by large language models on various reasoning tasks, models still suffer from severe drawbacks which shows a large gap with humans.

多跳推理大模型外部知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。