为小模型早期训练设计科学知识评估任务,提升训练阶段的可衡量性。
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
- 针对小模型早期训练设计专用评估任务,替代传统基准。
- 使用0.5B~3B参数模型,覆盖至200B词元训练的中间检查点。
- 支持免费云平台参与,适合非专家或资源有限的研究者。
现有基准在评估全量训练的大语言模型时表现良好,但在小模型早期训练阶段却难以提供有意义或具有区分度的信号。为探究此类差异的成因,本竞赛致力于设计专用于衡量语言模型早期训练进展的科学知识评估任务。参赛者需开发新型评估方法或改造现有基准,以更好捕捉不同模型间的性能差异。我们提供三个预训练小模型(0.5B、1B、3B参数),以及训练至200B词元期间的中间检查点。所有实验可在广泛可用的免费云GPU平台上运行,使计算资源有限的研究者也能参与。提交结果将根据性能信号质量、在1万亿词元训练后的模型排名一致性,以及与科学知识领域的相关性进行评估。该竞赛旨在推动早期训练阶段的评估策略创新,吸引跨学科参与者,促进大模型研究从训练初期就具备系统性和基准指导性。
原文摘要 · Abstract (English)
Existing benchmarks have proven effective for assessing the performance of fully trained large language models. However, we find striking differences in the early training stages of small models, where benchmarks often fail to provide meaningful or discriminative signals. To explore how these differences arise, this competition tackles the challenge of designing scientific knowledge evaluation tasks specifically tailored for measuring early training progress of language models. Participants are invited to develop novel evaluation methodologies or adapt existing benchmarks to better capture performance differences among language models. To support this effort, we provide three pre-trained small models (0.5B, 1B, and 3B parameters), along with intermediate checkpoints sampled during training up to 200B tokens. All experiments and development work can be run on widely available free cloud-based GPU platforms, making participation accessible to researchers with limited computational resources. Submissions will be evaluated based on three criteria: the quality of the performance signal they produce, the consistency of model rankings at 1 trillion tokens of training, and their relevance to the scientific knowledge domain. By promoting the design of tailored evaluation strategies for early training, this competition aims to attract a broad range of participants from various disciplines, including those who may not be machine learning experts or have access to dedicated GPU resources. Ultimately, this initiative seeks to make foundational LLM research more systematic and benchmark-informed from the earliest phases of model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。