指出大模型无法真正思考,并提出在特征空间中实现思考的架构改进方案。
Why LLMs Cannot Think and How to Fix It
- 定义'思考'并揭示当前大模型因架构限制无法产生真实思维
- 理论证明语言建模训练方式使模型无法在特征空间中进行思考
- 提出新架构设计,适合研究模型认知机制的学者参考
本文阐明,当前最先进的大型语言模型(LLMs)由于其架构约束,在特征空间中根本无法做出决策或产生‘思考’。我们建立了涵盖传统理解的‘思考’定义,并将其适配于大模型的应用场景。研究表明,当代大模型的架构设计与语言建模训练方法本质上排除了其进行真实思维过程的可能性。本文重点在于这一理论发现,而非实验数据带来的实践洞见。最后,我们提出了使特征空间内产生思维过程的解决方案,并讨论了这些架构修改带来的广泛影响。
原文摘要 · Abstract (English)
This paper elucidates that current state-of-the-art Large Language Models (LLMs) are fundamentally incapable of making decisions or developing "thoughts" within the feature space due to their architectural constraints. We establish a definition of "thought" that encompasses traditional understandings of that term and adapt it for application to LLMs. We demonstrate that the architectural design and language modeling training methodology of contemporary LLMs inherently preclude them from engaging in genuine thought processes. Our primary focus is on this theoretical realization rather than practical insights derived from experimental data. Finally, we propose solutions to enable thought processes within the feature space and discuss the broader implications of these architectural modifications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。