arXiv:2509.09009cs.LGcs.AI2025-09被引 6

开源1.7B参数模型,提供可复现的基准测试体系

Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison

  • 训练0.13B到1.7B参数模型,覆盖1T token规模
  • 在8个开放数据集上建立可比基准,验证训练有效性
  • 适合想对比训练方法、评估模型性能的研究者

我们提出 open-sci-ref,一组基于密集Transformer架构的基准模型,涵盖0.13B至1.7B参数规模及最高达1T token的训练量,基于8个近期公开参考数据集进行训练。在多个标准化基准上评估,这些训练结果建立了可比较的参考点,使研究者能检验不同训练方法在不同规模与数据集上的合理性与质量。中间检查点支持训练动态分析,通过统一计算资源轴对齐缩放趋势,实现训练流程的系统性对比。对比发现,在NemoTron-CC HQ上训练表现最优,其次为DCLM-baseline和FineWeb-Edu。除中间检查点外,发布还包含训练日志、代码及下游评估结果,以简化复现、统一比较,推动未来研究。

原文摘要 · Abstract (English)

We introduce open-sci-ref, a family of dense transformer models trained as research baselines across multiple model (0.13B to 1.7B parameters) and token scales (up to 1T) on 8 recent open reference datasets. Evaluating the models on various standardized benchmarks, our training runs set establishes reference points that enable researchers to assess the sanity and quality of alternative training approaches across scales and datasets. Intermediate checkpoints allow comparison and studying of the training dynamics. The established reference baselines allow training procedures to be compared through their scaling trends, aligning them on a common compute axis. Comparison of open reference datasets reveals that training on NemoTron-CC HQ consistently outperforms other reference datasets, followed by DCLM-baseline and FineWeb-Edu. In addition to intermediate training checkpoints, the release includes logs, code, and downstream evaluations to simplify reproduction, standardize comparison, and facilitate future research.

语言模型可复现基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。