arXiv:2603.16430cs.CLcs.AI2026-03

16B参数的开源大模型,高效且符合欧盟法规,支持多语言推理。

EngGPT2: Sovereign, Efficient and Open Intelligence

  • 从零训练的混合专家架构,每轮仅激活30亿参数,大幅降低算力需求。
  • 2.5万亿词训练,性能媲美8B-16B级稠密模型,推理功耗仅为1/5到1/2。
  • 专为欧洲和意大利语任务优化,支持中英双语推理与实时紧凑推理模式。

EngGPT2-16B-A3B是工程集团最新推出的意大利语大模型,旨在实现主权、高效与开放。该模型在2.5万亿词上训练,远少于Qwen3的36万亿或Llama3的15万亿,却在MMLU-Pro、GSM8K、IFEval和HumanEval等关键基准上达到8B-16B级稠密模型的性能水平,推理所需算力仅为后者的1/5至1/2,训练数据量和训练功耗也仅为1/10至1/6。采用从零训练的混合专家(MoE)架构,共160亿参数,每次推理仅激活30亿参数,专家规模介于GPT-OSS与Qwen3之间。约25%的训练语料为意大利语,显著提升欧洲及意大利语自然语言处理能力。模型支持非推理、意英双语推理及实时简洁推理模式(turbo-reasoning),并完全符合欧盟《人工智能法案》要求。其目标是为欧洲开源大模型生态树立资源节约型高性能新标准。

原文摘要 · Abstract (English)

EngGPT2-16B-A3B is the latest iteration of Engineering Group's Italian LLM and it's built to be a Sovereign, Efficient and Open model. EngGPT2 is trained on 2.5 trillion tokens - less than Qwen3's 36T or Llama3's 15T - and delivers performance on key benchmarks, including MMLU-Pro, GSM8K, IFEval and HumanEval, comparable to dense models in the 8B-16B range, while requiring one-fifth to half of the inference power, and between one-tenth to one-sixth of the training data and consequent needed training power. Designed as a trained-from-scratch Mixture-of-Experts (MoE) architecture, EngGPT2 features 16 billion parameters with 3 billion active per inference, with expert sizes positioned between those used in GPT-OSS and Qwen3. Approximately 25% of its training corpus consists of Italian-language data, to deliver strong capabilities for European and Italian NLP tasks among models of similar scale. This efficiency aims to position EngGPT2 as a key contributor to the growing portfolio of open-weight European models, combining performance and efficiency with full alignment to the EU AI Act. EngGPT2 is also a single model capable of multiple reasoning modes: non-reasoning, reasoning in Italian or English, and turbo-reasoning (a concise, bullet-point style reasoning available in both languages designed for real-time reasoning use cases). EngGPT2 aims to set a new standard for resource-conscious, high-performance LLMs tailored to European and Italian contexts.

大模型MoE开源欧洲AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。