arXiv:2601.11659cs.SEcs.LG2026-01被引 26

Llama 4技术细节全解析,涵盖架构、训练与部署要点。

The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes

  • 拆解Llama 4多专家混合架构及长上下文设计
  • 披露Scout与Maverick模型在基准测试中的性能表现
  • 适合关注大模型落地部署与合规使用的开发者

本文整合了Meta公开的Llama 4系列模型技术信息,涵盖发布版本(Scout与Maverick)及其整体生态背景,包括预览版的Behemoth教师模型。内容包括:超越高层MoE描述的架构细节,如路由/共享专家结构、早期融合多模态设计、以及Scout的长上下文特性(iRoPE与长度泛化策略);训练流程披露,包含预训练、中段长上下文扩展训练及后训练方法(轻量级SFT、在线RL与轻量级DPO);开发者报告的基座与指令微调版本基准结果;主流服务环境中的实际部署约束,包括供应商特定上下文长度限制和量化打包方式。此外,还总结了关于再分发与衍生命名的许可义务,以及公开描述的安全防护与评估实践。目标是为研究人员和从业者提供一份基于来源、精准可靠的技术参考。

原文摘要 · Abstract (English)

This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd context including the previewed Behemoth teacher model, (ii) architectural characteristics beyond a high-level MoE description covering routed/shared-expert structure, early-fusion multimodality, and long-context design elements reported for Scout (iRoPE and length generalization strategies), (iii) training disclosures spanning pre-training, mid-training for long-context extension, and post-training methodology (lightweight SFT, online RL, and lightweight DPO) as described in release materials, (iv) developer-reported benchmark results for both base and instruction-tuned checkpoints, and (v) practical deployment constraints observed across major serving environments, including provider-specific context limits and quantization packaging. The manuscript also summarizes licensing obligations relevant to redistribution and derivative naming, and reviews publicly described safeguards and evaluation practices. The goal is to provide a compact technical reference for researchers and practitioners who need precise, source-backed facts about Llama 4.

大模型架构分析部署优化开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。