arXiv:2601.07696cs.CL2026-01被引 2

用工具式多跳表格问答测试大模型的元推理能力。

Exploring the Meta-level Reasoning of Large Language Models via a Tool-based Multi-hop Tabular Question Answering Task

  • 设计基于地缘指标的多步问答任务,区分元推理与对象推理。
  • 模型能合理选工具但常犯数学错误,数值理解能力弱。
  • 适合研究大模型推理机制或评估其逻辑规划能力者阅读。

当前大语言模型的研究日益聚焦于‘推理’能力,但该概念在学术讨论中存在多重重叠定义。本文提出更结构化的区分:元级推理(关于完成任务所需中间步骤的思考)与对象级推理(执行这些步骤的低层操作)。为此,设计了一种新型问答任务,基于各国不同年份的地缘指标数据,问题需分解为多个中间步骤、检索数据并进行数学运算。通过分析模型选择合适工具的能力来评估其元级推理水平。为深入评估模型表现,任务引入‘必要操作’作为基准,对比模型工具调用输出以推断推理强度。结果表明,模型在该任务中展现出良好的元级推理能力,但在任务理解上仍存在缺陷。n-shot提示对准确率影响甚微;错误信息通常不会显著降低性能;进一步证实了模型在数值理解上的不足。最后讨论了研究发现对其他任务领域的泛化性与局限性。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) are increasingly focused on "reasoning" ability, a concept with many overlapping definitions in the LLM discourse. We take a more structured approach, distinguishing meta-level reasoning (denoting the process of reasoning about intermediate steps required to solve a task) from object-level reasoning (which concerns the low-level execution of the aforementioned steps.) We design a novel question answering task, which is based around the values of geopolitical indicators for various countries over various years. Questions require breaking down into intermediate steps, retrieval of data, and mathematical operations over that data. The meta-level reasoning ability of LLMs is analysed by examining the selection of appropriate tools for answering questions. To bring greater depth to the analysis of LLMs beyond final answer accuracy, our task contains 'essential actions' against which we can compare the tool call output of LLMs to infer the strength of reasoning ability. We find that LLMs demonstrate good meta-level reasoning on our task, yet are flawed in some aspects of task understanding. We find that n-shot prompting has little effect on accuracy; error messages encountered do not often deteriorate performance; and provide additional evidence for the poor numeracy of LLMs. Finally, we discuss the generalisation and limitation of our findings to other task domains.

大模型推理元推理表格问答工具使用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。