arXiv:2505.22296cs.CLcs.LG2025-05被引 2

360-LLaMA-Factory实现即插即用的序列并行,加速长文本微调。

360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training

  • 提出可直接接入LLaMA-Factory的序列并行方案。
  • 支持Light-R1、TinyR1等模型训练,提升长文本处理效率。
  • 适合需要高效微调大模型的科研与工业团队。

将序列并行技术集成至LLaMA-Factory,我们开源了360-LLaMA-Factory(https://github.com/Qihoo360/360-LLaMA-Factory)。该工具已获得广泛认可,被用于Light-R1(arXiv:2503.10460)、TinyR1(arXiv:2503.04872)、Kaggle AIMO数学模型,以及多家大型企业的训练框架中。本技术报告深入探讨了360-LLaMA-Factory背后的多种序列并行模式,并分享了关键实现经验。

原文摘要 · Abstract (English)

Adding sequence parallelism into LLaMA-Factory, we open-sourced 360-LLaMA-Factory at https://github.com/Qihoo360/360-LLaMA-Factory. 360-LLaMA-Factory has received wide recognition and used in models such as Light-R1 arXiv:2503.10460, TinyR1 arXiv:2503.04872, Kaggle AIMO math models and also in large companies' training frameworks. This technical report delves deeper into the different sequence parallel modes behind 360-LLaMA-Factory and discusses our implementation insights.

序列并行LLaMA微调加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。