探索大模型真正使用的算法,揭示其推理机制
Position: We Need An Algorithmic Understanding of Generative AI
- 提出AlgEval框架,系统分析模型隐层、注意力与计算过程中的算法特征
- 通过注意力模式与隐藏状态的电路级分析,验证搜索算法的涌现
- 为可解释性提供新路径,助力高效训练与新型架构设计
大型语言模型实际学习并使用哪些算法来解决问题?现有研究对此关注极少,因研究重心集中于通过规模提升性能,导致理论与实证层面存在理解空白。本文提出AlgEval:一个系统研究模型所学算法的框架,旨在揭示隐层表示、注意力机制与推理时计算中反映的算法基元及其组合方式,以解决特定任务。文章探讨了潜在的方法路径,并以涌现搜索算法为例展开案例研究,展示自上而下的算法假设形成与自下而上的注意力模式及隐藏状态的电路级验证。对模型如何真正解决问题的严谨系统评估,可替代资源密集型的规模扩展,推动领域转向对底层计算原理的理性理解。此类算法解释有助于实现人类可理解的可解释性,增进对模型内部推理的理解,进而促进更高效的训练方法与性能提升,以及端到端与多智能体系统的新型架构发展。
原文摘要 · Abstract (English)
What algorithms do LLMs actually learn and use to solve problems? Studies addressing this question are sparse, as research priorities are focused on improving performance through scale, leaving a theoretical and empirical gap in understanding emergent algorithms. This position paper proposes AlgEval: a framework for systematic research into the algorithms that LLMs learn and use. AlgEval aims to uncover algorithmic primitives, reflected in latent representations, attention, and inference-time compute, and their algorithmic composition to solve task-specific problems. We highlight potential methodological paths and a case study toward this goal, focusing on emergent search algorithms. Our case study illustrates both the formation of top-down hypotheses about candidate algorithms, and bottom-up tests of these hypotheses via circuit-level analysis of attention patterns and hidden states. The rigorous, systematic evaluation of how LLMs actually solve tasks provides an alternative to resource-intensive scaling, reorienting the field toward a principled understanding of underlying computations. Such algorithmic explanations offer a pathway to human-understandable interpretability, enabling comprehension of the model's internal reasoning performance measures. This can in turn lead to more sample-efficient methods for training and improving performance, as well as novel architectures for end-to-end and multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。