arXiv:2604.26752cs.CV2026-04被引 18

GLM-5V-Turbo让多模态感知成为智能体的核心能力

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

论文配图:GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
图 1 · 摘自论文原文
  • 将图像、视频等多模态输入深度融入推理与执行流程
  • 在多模态编程和视觉工具使用任务中表现优异
  • 适合构建真实场景下自主决策的多模态智能体

我们提出 GLM-5V-Turbo,迈向原生多模态智能体的基础模型。随着基础模型在真实环境中的部署增多,智能体的能力不仅依赖语言推理,还需具备对图像、视频、网页、文档、GUI等异构上下文的感知、理解与行动能力。GLM-5V-Turbo 的设计核心是将多模态感知作为推理、规划、工具调用与执行的内在组成部分,而非语言模型的附加接口。本报告总结了模型设计、多模态训练、强化学习、工具链扩展及智能体框架集成等方面的改进。这些进展使模型在多模态编程、视觉工具使用和基于框架的智能体任务中表现强劲,同时保持了优秀的纯文本编程能力。更重要的是,开发过程提供了构建多模态智能体的实用洞见,强调多模态感知的核心作用、分层优化策略以及可靠的端到端验证机制。

原文摘要 · Abstract (English)

We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability to perceive, interpret, and act over heterogeneous contexts such as images, videos, webpages, documents, GUIs. GLM-5V-Turbo is built around this objective: multimodal perception is integrated as a core component of reasoning, planning, tool use, and execution, rather than as an auxiliary interface to a language model. This report summarizes the main improvements behind GLM-5V-Turbo across model design, multimodal training, reinforcement learning, toolchain expansion, and integration with agent frameworks. These developments lead to strong performance in multimodal coding, visual tool use, and framework-based agentic tasks, while preserving competitive text-only coding capability. More importantly, our development process offers practical insights for building multimodal agents, highlighting the central role of multimodal perception, hierarchical optimization, and reliable end-to-end verification.

多模态智能体视觉推理工具使用基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。