跨平台测试15个大模型推理能力,发现数据质量比模型大小更重要。
Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- 在超算、云平台和高校集群上测试,验证结果可复现
- 79个学科问题中,训练数据质量影响远超模型规模
- 为教育、科研、生产提供选型参考,支持长期追踪
本文对当前基础模型的推理能力进行了跨平台全面评估,构建了覆盖三种计算范式(超算MareNostrum 5、云平台Nebius AI Studio、高校集群8张H200 GPU节点)的无基础设施依赖基准。评估涵盖15个基础模型,在8个学科领域(物理、数学、化学、经济、生物、统计、微积分、优化)的79个问题上展开三阶段实验:(1)基线建立:6个模型(Mixtral-8x7B、Phi-3、LLaMA 3.1-8B、Gemma-2-9b、Mistral-7B、OLMo-7B)在MareNostrum 5上测试19个问题,确立方法与基准;(2)基础设施验证:在高校集群(7个模型,含Falcon-Mamba)和Nebius平台(9个前沿模型:Hermes-4 70B/405B、LLaMA 3.1-405B/3.3-70B、Qwen3 30B/235B、DeepSeek-R1、GPT-OSS 20B/120B)重测19题,确认可复现性;(3)扩展评估:在高校集群与Nebius平台完成全部79题,探测架构多样性下的泛化能力。结果挑战传统规模假设,强调训练数据质量高于模型规模,并为教育、生产与研究场景提供可操作选型指南。三平台方法与79题基准支持大模型推理能力的长期追踪。
原文摘要 · Abstract (English)
This paper presents a comprehensive cross-platform evaluation of reasoning capabilities in contemporary foundation models, establishing an infrastructure-agnostic benchmark across three computational paradigms: HPC supercomputing (MareNostrum 5), cloud platforms (Nebius AI Studio), and university clusters (a node with eight H200 GPUs). We evaluate 15 foundation models across 79 problems spanning eight academic domains (Physics, Mathematics, Chemistry, Economics, Biology, Statistics, Calculus, and Optimization) through three experimental phases: (1) Baseline establishment: Six models (Mixtral-8x7B, Phi-3, LLaMA 3.1-8B, Gemma-2-9b, Mistral-7B, OLMo-7B) evaluated on 19 problems using MareNostrum 5, establishing methodology and reference performance; (2) Infrastructure validation: The 19-problem benchmark repeated on university cluster (seven models including Falcon-Mamba state-space architecture) and Nebius AI Studio (nine state-of-the-art models: Hermes-4 70B/405B, LLaMA 3.1-405B/3.3-70B, Qwen3 30B/235B, DeepSeek-R1, GPT-OSS 20B/120B) to confirm infrastructure-agnostic reproducibility; (3) Extended evaluation: Full 79-problem assessment on both university cluster and Nebius platforms, probing generalization at scale across architectural diversity. The findings challenge conventional scaling assumptions, establish training data quality as more critical than model size, and provide actionable guidelines for model selection across educational, production, and research contexts. The tri-infrastructure methodology and 79-problem benchmark enable longitudinal tracking of reasoning capabilities as foundation models evolve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。