系统梳理大模型推理能力突破,揭示未来研究关键挑战。
Reasoning Beyond Limits: Advances and Open Problems for LLMs
- 综述2023-2025年27个主流大模型的推理技术演进。
- 指出多步推理、跨语言能力与长上下文处理是核心进展。
- 适合关注大模型智能边界与未来方向的研究者阅读。
近年来生成式推理的突破彻底改变了大语言模型(LLMs)应对复杂任务的方式,使其能够动态检索、优化并组织信息形成连贯的多步推理链。包括DeepSeek-R1、OpenAI o1和o3、GPT-4o、Qwen-32B及多种Llama变体在内的先进模型,通过推理时扩展、强化学习、监督微调和知识蒸馏等技术显著提升了推理能力。本文全面回顾了2023至2025年间发布的27个主流大模型,如Mistral AI Small 3 24B、DeepSeek-R1、Search-o1、QwQ-32B和Phi-4,分析其核心创新与性能提升。同时,详细探讨了多语言大模型(MLLMs)在跨语言推理方面的进展,以及如何克服英语中心训练的局限性。此外,对基于状态空间模型(SSM)的架构(如Mamba)进行了系统综述,其在长上下文处理上相比传统Transformer更具效率。本文还覆盖了通用优化策略、专家混合(MoE)、检索增强生成(RAG)、思维链提示、自提升方法、测试时计算扩展与知识蒸馏框架等训练范式。最后,识别出未来研究的关键挑战:实现无需人工监督的多步推理、提升链式任务执行鲁棒性、平衡结构化提示与生成灵活性,以及加强长上下文检索与外部工具的集成。
原文摘要 · Abstract (English)
Recent breakthroughs in generative reasoning have fundamentally reshaped how large language models (LLMs) address complex tasks, enabling them to dynamically retrieve, refine, and organize information into coherent multi-step reasoning chains. Techniques such as inference-time scaling, reinforcement learning, supervised fine-tuning, and distillation have been effectively applied to state-of-the-art models, including DeepSeek-R1, OpenAI o1 and o3, GPT-4o, Qwen-32B, and various Llama variants, significantly enhancing their reasoning capabilities. In this paper, we present a comprehensive review of the top 27 LLMs released between 2023 and 2025, such as Mistral AI Small 3 24B, DeepSeek-R1, Search-o1, QwQ-32B, and Phi-4, and analyze their core innovations and performance improvements. We also provide a detailed overview of recent advancements in multilingual large language models (MLLMs), emphasizing methods that improve cross-lingual reasoning and address the limitations of English-centric training. In parallel, we present a comprehensive review of progress in state space model (SSM)-based architectures, including models such as Mamba, which demonstrate improved efficiency for long-context processing compared to transformer-based approaches. Our analysis covers training strategies including general optimization techniques, mixture-of-experts (MoE) configurations, retrieval-augmented generation (RAG), chain-of-thought prompting, self-improvement methods, and test-time compute scaling and distillation frameworks. Finally, we identify key challenges for future research, including enabling multi-step reasoning without human supervision, improving robustness in chained task execution, balancing structured prompting with generative flexibility, and enhancing the integration of long-context retrieval and external tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。