arXiv:2503.11314cs.CL2025-03ACL被引 37

通过表征工程激发大模型通用长链推理能力

Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation Engineering

  • 从表征角度分析,发现大模型具备通用长链推理能力
  • 仅需少量示例微调,即可在跨领域任务中有效迁移
  • 适合希望提升模型推理泛化能力的研究者

近期长链思维(long CoT)进展显著提升了大语言模型(LLMs)的推理能力。现有研究发现,仅需少量示例微调即可高效激发长链推理能力,并可轻松迁移到其他任务。这促使我们探究长链推理是否为大模型的通用能力。本文从表征视角进行实证分析,发现大模型确实编码了长链推理的通用能力,且与普通链式思维有明显区别。此外,领域特定表征对长链推理的有效迁移至关重要。基于此,我们提出 GLoRE,一种新型表征工程方法,以释放大模型的通用长链推理潜力。大量实验表明,GLoRE在同域和跨域场景下均具有高效性与有效性。

原文摘要 · Abstract (English)

Recent advancements in long chain-of-thoughts(long CoTs) have significantly improved the reasoning capabilities of large language models(LLMs). Existing work finds that the capability of long CoT reasoning can be efficiently elicited by tuning on only a few examples and can easily transfer to other tasks. This motivates us to investigate whether long CoT reasoning is a general capability for LLMs. In this work, we conduct an empirical analysis for this question from the perspective of representation. We find that LLMs do encode long CoT reasoning as a general capability, with a clear distinction from vanilla CoTs. Furthermore, domain-specific representations are also required for the effective transfer of long CoT reasoning. Inspired by these findings, we propose GLoRE, a novel representation engineering method to unleash the general long CoT reasoning capabilities of LLMs. Extensive experiments demonstrate the effectiveness and efficiency of GLoRE in both in-domain and cross-domain scenarios.

长链推理表征工程大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。