用电路模型解释大模型推理机制,提升预测准确率。
Electronic Circuit Principles of Large Language Models
- 将大模型推理映射为电路网络,基于欧姆与法拉第定律建模。
- 在350个任务上相关性提升60%,优于传统缩放定律。
- 可解释15种提示策略,指导设计更优模块化干预方法。
大型语言模型(如DeepSeek-R1)在多种推理任务中表现卓越。为揭示其行为规律,本文提出电子电路原理(ECP),将推理时学习(ITL)映射为语义电动势,推理时推理(ITR)映射为受欧姆与法拉第定律支配的电阻网络。该电路模型实现了任务性能的闭式预测,并揭示了模块化提示组件如何协同影响准确率。我们在70,000个样本、350个推理任务和9个先进LLM上验证了ECP,观察到其皮尔逊相关性相比传统推理时缩放定律提升约60%。此外,ECP解释了15种已知提示策略的有效性,并指导开发出的新模块化干预方法,在国际信息学奥林匹克和国际数学奥林匹克竞赛中均超过前80%参赛者的平均分。通过将大模型推理建立在电子电路原理基础上,ECP提供了一个严谨的性能预测与模块优化框架。
原文摘要 · Abstract (English)
Large language models (LLMs) such as DeepSeek-R1 have achieved remarkable performance across diverse reasoning tasks. To uncover the principles that govern their behaviour, we introduce the Electronic Circuit Principles (ECP), which maps inference-time learning (ITL) onto a semantic electromotive force and inference-time reasoning (ITR) onto a resistive network governed by Ohm's and Faraday's laws. This circuit-based modelling yields closed-form predictions of task performance and reveals how modular prompt components interact to shape accuracy. We validated ECP on 70,000 samples spanning 350 reasoning tasks and 9 advanced LLMs, observing a about 60% improvement in Pearson correlation relative to the conventional inference-time scaling law. Moreover, ECP explains the efficacy of 15 established prompting strategies and directs the development of new modular interventions that exceed the median score of the top 80% of participants in both the International Olympiad in Informatics and the International Mathematical Olympiad. By grounding LLM reasoning in electronic-circuit principles, ECP provides a rigorous framework for predicting performance and optimising modular components.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。