arXiv:2507.22457cs.CLcs.AI2025-07被引 2

大模型看似不擅抽象推理,但微调部分参数即可大幅提升表现。

What is an "Abstract Reasoner"? Revisiting Experiments and Arguments about Large Language Models

  • 通过微调输入编码参数,零样本差的表现可接近完美
  • 微调效果无法跨数据集迁移,表明泛化能力有限
  • 挑战了大模型是否为抽象推理者的定义,适合思考模型本质的研究者

近期研究认为大语言模型(LLMs)并非'抽象推理者',依据是其在多种挑战性任务中零样本表现不佳。本文重新审视这些实验,发现尽管LLMs在零样本设置下表现较差,但仅对少量输入编码参数进行微调即可实现近乎完美的性能。然而,这种微调效果并不具备跨数据集的泛化能力。基于这一系列实证结果,本文呼吁重新讨论'抽象推理者'的定义及其重要性,以及大模型是否符合该标准。

原文摘要 · Abstract (English)

Recent work has argued that large language models (LLMs) are not "abstract reasoners", citing their poor zero-shot performance on a variety of challenging tasks as evidence. We revisit these experiments in order to add nuance to the claim. First, we show that while LLMs indeed perform poorly in a zero-shot setting, even tuning a small subset of parameters for input encoding can enable near-perfect performance. However, we also show that this finetuning does not necessarily transfer across datasets. We take this collection of empirical results as an invitation to (re-)open the discussion of what it means to be an "abstract reasoner", and why it matters whether LLMs fit the bill.

大模型推理能力微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。