arXiv:2411.11371cs.CLcs.LG2024-11被引 2

思考标记在实践中表现不佳,因单一嵌入导致学习信号不一致。

Rethinking Thinking Tokens: Understanding Why They Underperform in Practice

  • 用单一嵌入表示思考过程,引发学习信号波动
  • 在多个基准上均弱于思维链推理,提升微乎其微
  • 适合研究无监督推理机制的学者参考

思考标记(Thinking Tokens, TT)被提出作为语言模型中促进推理的无监督方法。然而,尽管概念上有吸引力,我们的研究发现,TT仅带来微小性能提升,且在多个基准测试中始终弱于思维链(Chain-of-Thought, CoT)推理。我们推测,这种表现不佳源于TT依赖单一嵌入,导致学习信号不一致并引入噪声梯度。本文通过全面的实证分析验证了该假设,并讨论了对未来大模型无监督推理研究的启示。

原文摘要 · Abstract (English)

Thinking Tokens (TT) have been proposed as an unsupervised method to facilitate reasoning in language models. However, despite their conceptual appeal, our findings show that TTs marginally improves performance and consistently underperforms compared to Chain-of-Thought (CoT) reasoning across multiple benchmarks. We hypothesize that this underperformance stems from the reliance on a single embedding for TTs, which results in inconsistent learning signals and introduces noisy gradients. This paper provides a comprehensive empirical analysis to validate this hypothesis and discusses the implications for future research on unsupervised reasoning in LLMs.

无监督推理思考标记语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。