剖析大模型推理机制,揭示训练、推理与错误背后的内在原理
Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures
- 从训练动态、推理机制、意外行为三方面系统梳理大模型内部运作
- 提出需突破黑箱性能,实现机制层面的透明理解
- 适合关注模型可解释性与可靠性研究的学者与工程师
强化学习(RL)推动了大型推理模型(LRMs)的发展,使其推理能力达到新高度。尽管性能令人振奋,但探究其内部行为机制已成为同样重要的研究前沿。本文全面综述了LRMs的机制理解,将近期成果归纳为三个核心维度:1)训练动态,2)推理机制,3)非预期行为。通过整合这些洞见,旨在弥合黑箱性能与机制透明性之间的差距。最后,我们讨论了未充分探索的挑战,提出了未来机制研究的路线图,包括应用可解释性需求、方法改进以及统一理论框架的建立。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has catalyzed the emergence of Large Reasoning Models (LRMs) that have pushed reasoning capabilities to new heights. While their performance has garnered significant excitement, exploring the internal mechanisms driving these behaviors has become an equally critical research frontier. This paper provides a comprehensive survey of the mechanistic understanding of LRMs, organizing recent findings into three core dimensions: 1) training dynamics, 2) reasoning mechanisms, and 3) unintended behaviors. By synthesizing these insights, we aim to bridge the gap between black-box performance and mechanistic transparency. Finally, we discuss under-explored challenges to outline a roadmap for future mechanistic studies, including the need for applied interpretability, improved methodologies, and a unified theoretical framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。