arXiv:2512.14427cs.CLcs.AI2025-12

文档打包能提升大模型多跳推理能力,但需更多算力。

Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models

  • 通过不同文档打包策略测试模型推理能力。
  • 打包后模型性能提升,但训练耗时增加。
  • 揭示了打包提升效果的关键机制,适合模型优化研究者。

大语言模型的标准训练方法是将多个文档打包以提升计算效率,但这一过程对模型能力的影响仍不明确。本文研究不同文档打包策略对大模型潜在多跳推理能力的影响。结果表明,相比单独训练文档,打包可提升模型性能,但需要更多计算资源。通过消融实验,我们识别出影响打包优势的关键因素。本研究深化了对大模型训练机制的理解,为优化模型开发提供了实用洞见。

原文摘要 · Abstract (English)

The standard practice for training large language models involves packing multiple documents together to optimize computational efficiency. However, the impact of this process on the models' capabilities remains largely unexplored. To address this gap, we investigate how different document-packing strategies influence the latent multi-hop reasoning abilities of LLMs. Our findings indicate that packing can improve model performance compared to training on individual documents, at the expense of more compute. To further understand the underlying mechanisms, we conduct an ablation study, identifying key factors that explain the advantages of packing. Ultimately, our research deepens the understanding of LLM training dynamics and provides practical insights for optimizing model development.

大模型训练多跳推理文档打包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。