arXiv:2502.10338cs.CLcs.AI2025-02中稿 · the Workshop on Pl…被引 1

测试大模型在复杂问答中的元层次与对象层次推理能力。

Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering

  • 区分元层次战略推理与对象层次具体计算
  • 大模型在元层次推理中表现良好,对象层次较弱
  • 新数据集Franklin凸显模型在具体推理上的短板

大型语言模型在自然语言任务中表现优异,但在需要多步复杂推理的问答任务中仍存在挑战。本文梳理了此类任务所需的推理类型,将其分为元层次推理(类似高层策略规划)和对象层次推理(如数学推导等低层任务)。为此构建了新数据集Franklin,结合三个现有数据集,评估四种大模型在多步推理问答任务中的表现。人工标注结果显示,大模型在元层次推理中频繁出现,但对部分数据集中的对象层次推理任务表现不佳。尤其在Franklin数据集中,模型虽具备较强的元层次推理能力,但面对其中要求的对象层次推理仍显吃力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel in natural language tasks but still face challenges in Question Answering (QA) tasks requiring complex, multi-step reasoning. We outline the types of reasoning required in some of these tasks, and reframe them in terms of meta-level reasoning (akin to high-level strategic reasoning or planning) and object-level reasoning (embodied in lower-level tasks such as mathematical reasoning). Franklin, a novel dataset with requirements of meta- and object-level reasoning, is introduced and used along with three other datasets to evaluate four LLMs at question answering tasks requiring multiple steps of reasoning. Results from human annotation studies suggest LLMs demonstrate meta-level reasoning with high frequency, but struggle with object-level reasoning tasks in some of the datasets used. Additionally, evidence suggests that LLMs find the object-level reasoning required for the questions in the Franklin dataset challenging, yet they do exhibit strong performance with respect to the meta-level reasoning requirements.

大模型推理能力问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。