16B参数的意大利工程模型在多项国际评测中表现优异,优于多数本土同类模型。
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs

- 采用混合专家架构,仅激活30亿参数实现高效推理。
- 在多国基准测试中超越多数意大利开源模型,32K上下文长序列表现最佳。
- 适合关注本地化大模型性能与工业级应用的开发者和研究者。
本报告对意大利工程公司Ingegneria Informatica S.p.A.研发的EngGPT2MoE-16B-A3B模型进行评测,该模型为160亿参数的混合专家(MoE)结构,仅激活30亿参数。评测覆盖多种代表性基准,与同规模开源的MoE及密集型模型对比。在国际基准如ARC-Challenge、GSM8K、AIME24、AIME25、MMLU、HumanEval(HE)上表现达或优于意大利主流模型(FastwebMIIA-7B、Minerva-7B、Velvet-14B、LLaMAntino-3-ANITA-8B)。在32k上下文设置的RULER基准中表现最优。在意大利语数据集ITALIC上优于除Velvet-14B外的所有对比模型。相比同规模MoE模型,其在大多数基准上优于DeepSeek-MoE-16B-Chat;在HE、MMLU、AIME24、AIME25、GSM8K及32k RULER上优于Moonlight-16B-A3B,但在BFCL及部分ARC和ITALIC设置上略低。低于GPT-OSS-20B在多数任务上的表现。与密集模型对比,在AIME24/AIME25上优于Llama-3.1-8B-Instruct、Gemma-3-12b-it、Ministral-3-8BInstruct-2512-BF16,但在ITALIC、BFCL和32k RULER上表现较弱。综合所有指标,其性能优于所评意大利模型,但不及部分顶尖国际模型(如GPT-5 nano、Qwen3-8B)。总体表明该模型是意大利本土大模型的重要进展。
原文摘要 · Abstract (English)
This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE) model with 3B active parameters. Performance is investigated across a wide variety of representative benchmarks, and is compared against comparably-sized open-source MoE and dense models. In comparison with popular Italian models, namely FastwebMIIA-7B, Minerva-7B, Velvet-14B, and LLaMAntino-3-ANITA-8B, EngGPT2MoE-16B-A3B performs as well or better on international benchmarks: ARC-Challenge, GSM8K, AIME24, AIME25, MMLU, and HumanEval (HE). It achieves the best performance for the longest context setting (32k) of the RULER benchmark. On the Italian benchmark dataset ITALIC, the model performs as well or better than the other models except for Velvet-14B, which outperforms it. Compared with popular MoE models of comparable size, the new model reports higher values than DeepSeek-MoE-16B-Chat on all considered benchmarks. It has higher values than Moonlight-16B-A3B on HE, MMLU, AIME24, AIME25, GSM8K, and the 32k RULER setting, but lower on BFCL and some ARC and ITALIC settings. Finally it has lower values than GPT-OSS-20B on most benchmarks, including HE, MMLU, AIME24, AIME25, GSM8K, ARC, BFCL, and the RULER 32k. When compared with popular dense models, EngGPT2MoE-16B-A3B reports higher values on AIME24 and AIME25 than Llama-3.1-8B-Instruct, Gemma-3-12b-it, and Ministral-3-8BInstruct-2512-BF16, but lower values on ITALIC, BFCL, and RULER with a 32k context. When performance is aggregated across all benchmark metrics, EngGPT2MoE-16B-A3B shows higher performance than the Italian models under evaluation while achieving lower results than some of the most performant international models, in particular GPT-5 nano and Qwen3-8B. Taken together, our findings find the new model to be a step forward for native Italian Large Language Models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。