语言模型八年飞跃,性能暴涨成本骤降,专用模型成新趋势。
From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models

- 从BERT到智能代理,模型能力与效率同步提升
- 2024年后代码解决能力年增近6倍,顶尖模型仅需1-6美元/百万词元
- 细分任务模型崛起,如前端、仓库级和终端任务各有专长
2018年10月至2026年7月间,人工智能模型从BERT等简单系统演进为能解决复杂数学问题并编写软件的大型智能体。自2024年底以来,实际代码问题解决能力每年提升近六倍。与此同时,成本显著下降:OpenAI的预算型模型GPT 5.6 Luna以1至6美元每百万词元的价格,实现了旗舰级性能,远低于旧版本。当前顶级表现分散于专用模型——Claude Opus 5在前端编码领先,Claude Fable 5在仓库级编码占优,GPT 5.6 Sol则主导终端任务。在小学数学测试中,使用Qwen 2.5模型时,基础方法可解58道题,先进采样法最高达79道;置信度排序工具在前50项预测中准确识别出47个正确答案,对任务筛选极具价值,所有研究材料均已公开。
原文摘要 · Abstract (English)
Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software. The ability to resolve real coding issues improved by nearly six times per year since late 2024. During this time costs dropped sharply with OpenAIs budget model GPT 5 point 6 Luna matching flagship capabilities for just one to six dollars per million tokens beating older versions at a fraction of the price. Top performance is now split across specialized models as Claude Opus 5 leads in frontend coding Claude Fable 5 excels at repository level coding and GPT 5 point 6 Sol dominates terminal tasks. In a grade school math test using the Qwen 2 point 5 model basic methods solved 58 of 100 problems while advanced sampling solved up to 79. A confidence ranking tool correctly identified 47 right answers in its top 50 choices proving highly useful for sorting tasks with all research materials made fully public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。