发现机器学习在串行任务上存在根本局限,提出需重新设计模型与硬件。
The Serial Scaling Hypothesis
- 区分可并行与本质串行问题,建立复杂性理论框架。
- 首次证明扩散模型无法解决本质串行任务。
- 适合关注模型架构、硬件设计的从业者参考。
尽管机器学习通过大规模并行化取得进展,但我们识别出一个关键盲点:某些问题本质上是串行的。这类‘本质串行’问题——从数学推理到物理模拟再到序列决策——需要依赖前后步骤的计算依赖,无法高效并行化。我们在复杂性理论中形式化这一区分,并证明当前以并行为中心的架构在这些任务上存在根本限制。随后,我们首次展示扩散模型尽管具有串行特性,仍无法解决本质串行问题。我们认为,认识计算的串行本质对机器学习、模型设计和硬件发展具有深远影响。
原文摘要 · Abstract (English)
While machine learning has advanced through massive parallelization, we identify a critical blind spot: some problems are fundamentally sequential. These "inherently serial" problems-from mathematical reasoning to physical simulations to sequential decision-making-require sequentially dependent computational steps that cannot be efficiently parallelized. We formalize this distinction in complexity theory, and demonstrate that current parallel-centric architectures face fundamental limitations on such tasks. Then, we show for first time that diffusion models despite their sequential nature are incapable of solving inherently serial problems. We argue that recognizing the serial nature of computation holds profound implications on machine learning, model design, and hardware development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。