大模型在多跳推理中仍难有效利用外部知识
Large Language Models Still Face Challenges in Multi-Hop Reasoning with External Knowledge
- 测试GPT-3.5在四种链式思考提示下的多跳推理能力
- 模型在多跳任务中表现远低于人类,尤其在复杂推理时
- 适合研究大模型推理缺陷与提升方法的学者
我们通过一系列实验,从三个方面评估大语言模型的多跳推理能力:选择并组合外部知识、处理非顺序推理任务,以及推广到更多跳数的数据样本。我们在四个推理基准上测试了GPT-3.5模型,采用链式思考提示(Chain-of-Thought prompting)及其变体。结果表明,尽管大语言模型在多种推理任务中表现优异,但在多跳推理方面仍存在严重缺陷,与人类表现差距显著。
原文摘要 · Abstract (English)
We carry out a series of experiments to test large language models' multi-hop reasoning ability from three aspects: selecting and combining external knowledge, dealing with non-sequential reasoning tasks and generalising to data samples with larger numbers of hops. We test the GPT-3.5 model on four reasoning benchmarks with Chain-of-Thought prompting (and its variations). Our results reveal that despite the amazing performance achieved by large language models on various reasoning tasks, models still suffer from severe drawbacks which shows a large gap with humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。