arXiv:2411.00369cs.CL2024-11被引 4

构建首个带推理图结构的多跳问答数据集,揭示模型在不同推理路径上的表现差异。

GRS-QA -- Graph Reasoning-Structured Question Answering Dataset

  • 用推理图显式建模多跳问答中的逻辑路径,节点为文本片段,边表示推理关系。
  • 实验证明大模型在不同推理结构上表现差异显著,揭示结构影响性能。
  • 适合研究模型推理机制、评估推理能力或设计结构化数据集的研究者。

大型语言模型(LLMs)在多跳问答(M-QA)中表现出色,归因于其强大的推理能力。然而,内在推理结构对模型表现的影响仍不清晰,主要由于缺乏提供细粒度推理结构的问答数据集。为此,我们提出了图推理结构化问答数据集(GRS-QA),包含问答对的语义上下文和推理结构。与现有M-QA数据集不同,GRS-QA通过构建推理图显式捕捉复杂的推理路径:节点代表文本上下文,边表示逻辑流转。不同结构的推理图使我们能够对大模型在多种推理结构下的能力进行细粒度评估。实证分析显示,模型在处理不同推理结构的问题时表现各异,这一发现有助于探索文本结构与语义之间的关系。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have excelled in multi-hop question-answering (M-QA) due to their advanced reasoning abilities. However, the impact of the inherent reasoning structures on LLM M-QA performance remains unclear, largely due to the absence of QA datasets that provide fine-grained reasoning structures. To address this gap, we introduce the Graph Reasoning-Structured Question Answering Dataset (GRS-QA), which includes both semantic contexts and reasoning structures for QA pairs. Unlike existing M-QA datasets, where different reasoning structures are entangled together, GRS-QA explicitly captures intricate reasoning pathways by constructing reasoning graphs, where nodes represent textual contexts and edges denote logical flows. These reasoning graphs of different structures enable a fine-grained evaluation of LLM reasoning capabilities across various reasoning structures. Our empirical analysis reveals that LLMs perform differently when handling questions with varying reasoning structures. This finding facilitates the exploration of textual structures as compared with semantics.

多跳问答推理结构数据集大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。