让大模型在临界态推理,靠全局参数就能量化其推理能力。
PLDR-LLMs Reason At Self-Organized Criticality
- 在自组织临界态预训练,推理时输出呈现类相变特征。
- 临界态下相关长度发散,输出达稳定态,具备泛化与推理能力。
- 只需看推理时参数统计量,无需测试集即可评估推理水平。
我们发现,在自组织临界态预训练的PLDR-LLM在推理时表现出推理能力。临界态下,PLDR-LLM演绎输出的特性类似于二阶相变。在临界态,相关长度发散,演绎输出达到一种亚稳态稳定状态。该稳态行为表明,演绎输出学习到了与尺度函数、普适类和重整化群等价的表示,从而在过程中获得泛化与推理能力。我们可从推理时模型演绎输出参数的全局统计量定义一个序参量。当序参量接近零时,PLDR-LLM的推理能力更强。这一观察得到在近临界态与亚临界态训练的模型基准分数的支持。我们的结果为大语言模型中推理如何显现提供了自洽解释,且推理能力可仅通过推理时稳态下全局模型参数值来量化,无需通过归纳输出对精心构建的基准数据集进行评估。
原文摘要 · Abstract (English)
We show that PLDR-LLMs pretrained at self-organized criticality exhibit reasoning at inference time. The characteristics of PLDR-LLM deductive outputs at criticality is similar to second-order phase transitions. At criticality, the correlation length diverges, and the deductive outputs attain a metastable steady state. The steady state behaviour suggests that deductive outputs learn representations equivalent to scaling functions, universality classes and renormalization groups from the training dataset, leading to generalization and reasoning capabilities in the process. We can then define an order parameter from the global statistics of the model's deductive output parameters at inference. The reasoning capabilities of a PLDR-LLM is better when its order parameter is close to zero at criticality. This observation is supported by the benchmark scores of the models trained at near-criticality and sub-criticality. Our results provide a self-contained explanation on how reasoning manifests in large language models, and the ability to reason can be quantified solely from global model parameter values of the deductive outputs at steady state, without any need for evaluation of curated benchmark datasets through inductive output for reasoning and comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。