arXiv:2508.06471cs.CL2025-08被引 440

GLM-4.5 是一款高效推理与智能体系统模型,支持思考与直接回答双模式。

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

  • 采用混合专家架构,激活参数仅320亿,实现高效推理
  • 在TAU-Bench、AIME 24等任务上分别取得70.1%、91.0%得分
  • 开源大版与轻量版,适合研究推理与智能体系统

我们提出 GLM-4.5,一款开源的混合专家(MoE)大语言模型,总参数量达3550亿,激活参数为320亿,具备支持思考与直接响应的混合推理能力。通过在23万亿标记上进行多阶段训练,并结合专家模型迭代与强化学习的全面后训练,GLM-4.5在智能体、推理和编码(ARC)任务中表现优异,在TAU-Bench上得分为70.1%,AIME 24为91.0%,SWE-bench Verified为64.2%。相比多个竞品,其参数更少但排名位居前列——整体第三,智能体类任务第二。我们同时发布完整版GLM-4.5(355B)与轻量版GLM-4.5-Air(106B),以推动推理与智能体系统的研究。代码与模型详见 https://github.com/zai-org/GLM-4.5。

原文摘要 · Abstract (English)

We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes. Through multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning, GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks. We release both GLM-4.5 (355B parameters) and a compact version, GLM-4.5-Air (106B parameters), to advance research in reasoning and agentic AI systems. Code, models, and more information are available at https://github.com/zai-org/GLM-4.5.

大模型推理智能体开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。