梳理大模型发展脉络,揭示突破算力瓶颈的六大新范式。
LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems
- 构建循环式分类框架,覆盖2019-2025年50+模型与15个机构
- 发现数据、成本、能耗三重危机,提出6类破壁技术方案
- 聚焦效率革命与开源崛起,适配研究者与技术决策者
人工智能领域从基础Transformer架构演进至具备推理能力、逼近人类水平的系统。本文提出LLMOrbit,一个涵盖2019-2025年的大型语言模型综合循环分类体系,分析了来自15家机构的50余个模型,通过八个相互关联的轨道维度,记录现代大模型、生成式AI与智能体系统在架构创新、训练方法与效率模式上的演进。研究揭示三大危机:(1)数据稀缺(9–27万亿令牌将于2026–2028年耗尽),(2)成本指数增长(五年内从300万美元增至3亿美元以上),(3)能源消耗不可持续(增长22倍),确立了限制暴力扩展的“规模墙”。分析识别出六种突破该墙的范式:(1)推理时计算(o1、DeepSeek-R1以10倍推理算力达GPT-4性能),(2)量化压缩(4–8倍),(3)分布式边缘计算(成本降低10倍),(4)模型融合,(5)高效训练(ORPO减少50%内存),(6)小型专用模型(Phi-4 14B媲美更大模型)。三个范式跃迁浮现:(1)后训练增益(RLHF、GRPO、纯强化学习贡献显著,DeepSeek-R1在MATH上达79.8%),(2)效率革命(MoE路由提升18倍效率,多头潜在注意力实现8倍键值缓存压缩,使GPT-4级性能成本低于0.30美元/千令牌),(3)民主化(开源Llama 3在MMLU上达88.6%,超越GPT-4的86.4%)。提供对强化学习微调(RLHF、PPO、DPO、GRPO、ORPO)等技术的洞察,追踪从被动生成到工具使用智能体(ReAct、RAG、多智能体系统)的演变,并分析后训练创新。
原文摘要 · Abstract (English)
The field of artificial intelligence has undergone a revolution from foundational Transformer architectures to reasoning-capable systems approaching human-level performance. We present LLMOrbit, a comprehensive circular taxonomy navigating the landscape of large language models spanning 2019-2025. This survey examines over 50 models across 15 organizations through eight interconnected orbital dimensions, documenting architectural innovations, training methodologies, and efficiency patterns defining modern LLMs, generative AI, and agentic systems. We identify three critical crises: (1) data scarcity (9-27T tokens depleted by 2026-2028), (2) exponential cost growth ($3M to $300M+ in 5 years), and (3) unsustainable energy consumption (22x increase), establishing the scaling wall limiting brute-force approaches. Our analysis reveals six paradigms breaking this wall: (1) test-time compute (o1, DeepSeek-R1 achieve GPT-4 performance with 10x inference compute), (2) quantization (4-8x compression), (3) distributed edge computing (10x cost reduction), (4) model merging, (5) efficient training (ORPO reduces memory 50%), and (6) small specialized models (Phi-4 14B matches larger models). Three paradigm shifts emerge: (1) post-training gains (RLHF, GRPO, pure RL contribute substantially, DeepSeek-R1 achieving 79.8% MATH), (2) efficiency revolution (MoE routing 18x efficiency, Multi-head Latent Attention 8x KV cache compression enables GPT-4-level performance at $<$$0.30/M tokens), and (3) democratization (open-source Llama 3 88.6% MMLU surpasses GPT-4 86.4%). We provide insights into techniques (RLHF, PPO, DPO, GRPO, ORPO), trace evolution from passive generation to tool-using agents (ReAct, RAG, multi-agent systems), and analyze post-training innovations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。