arXiv:2607.02770cs.CLcs.AI2026-07被引 60

Gemma 4 推出多模态开源大模型,提升推理与效率。

Gemma 4 Technical Report

论文配图:Gemma 4 Technical Report
图 1 · 摘自论文原文
  • 采用密集与专家混合架构,参数量2.3B至31B,支持多模态输入。
  • 12B模型无需编码器,直接处理原始音视频块,推理速度更快。
  • 新增思考模式生成推理过程,长文本与跨领域任务表现领先。

我们推出新一代开源多模态语言模型 Gemma 4,属于 Gemma 模型家族。该系列模型在计算效率和推理能力方面实现显著提升,包含从 2.3B 到 31B 参数的密集型与混合专家(Mixture-of-Experts)架构。所有模型均配备优化后的视觉与音频编码器。针对 12B 模型,提出一种无编码器统一架构,可直接接收原始音频与图像块输入。此外,引入“思考模式”,使模型可在回复前生成推理轨迹。通过关键设计优化,显著提升推理速度、内存占用与计算效率,并增强长上下文处理能力。Gemma 4 在 STEM、多模态及长文本基准测试中表现跃升,人类评估任务中媲美更大规模的前沿开源模型。

原文摘要 · Abstract (English)

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

多模态开源模型推理优化Gemma

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。