arXiv:2511.21613cs.CLcs.AI2025-11被引 1

用细粒度元数据加速大模型预训练,提升效率与效果

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining

  • 将细粒度文档质量指标作为元数据前置,可加速预训练
  • 引入元数据附加机制,辅助任务预测提升训练速度
  • 可学习的元标记能恢复部分加速效果,适合优化训练流程

在大语言模型(LLM)预训练中引入元数据已成为加速训练的有前途方法。然而,以往研究仅关注URL这一单一信号,未探索其他元数据类型是否更具价值。本文系统考察了多种元数据类型,发现细粒度文档质量指标等信息在前置时也能显著加速预训练。我们识别出有效元数据的共同特征:编码信息粒度更细。进一步提出元数据附加策略,通过预测元数据作为辅助任务,提升训练效率。此外,使用掩码损失训练可学习元标记,能恢复部分加速效果,诱导出与质量相关的潜在结构。通过探针分析,揭示了元数据如何影响模型隐层表示。上述成果为提升LLM预训练效率与有效性提供了实用指导。

原文摘要 · Abstract (English)

Incorporating metadata in Large Language Models (LLMs) pretraining has recently emerged as a promising approach to accelerate training. However prior work highlighted only one useful signal-URLs, leaving open the question of whether other forms of metadata could yield greater benefits. In this study, we investigate a wider range of metadata types and find other types of metadata, such as fine-grained indicators of document quality that can also accelerate pretraining when prepended. We identify a common feature among effective metadata: they encode information at a finer granularity. We further introduce metadata appending as a means of improving training efficiency, where predicting an appropriate metadata as auxiliary task can help speed up pretraining. In addition, learnable meta-tokens trained with masked loss can recover part of the speedup by inducing quality-aware latent structure. Using probing, we analyze latent representations to understand how metadata shapes learning. Together, these results yield practical guidelines for integrating metadata to improve both the efficiency and effectiveness of LLM pretraining.

元数据预训练效率优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。