Ruyi2通过家族化参数共享实现高效变深度推理,训练一次可部署多版本。
Ruyi2 Technical Report
- 采用基于Megatron-LM的家族化模型架构,支持3D并行训练
- 相比Ruyi提升2-3倍速度,性能媲美同规模Qwen3模型
- 适用于需多版本部署的高并发场景,适合追求效率与性能平衡的研究者
大型语言模型在部署成本和延迟方面面临严峻挑战,亟需自适应计算策略。在AI Flow框架基础上,我们提出Ruyi2作为自适应模型系列的演进版本,旨在实现高效的变深度计算。尽管早期退出架构能较好平衡效率与性能,但现有Ruyi模型及方法常受限于优化复杂度,且难以适配大规模分布式训练。为解决此问题,Ruyi2引入基于Megatron-LM的稳定“家族化模型”,通过3D并行训练,相较Ruyi实现2-3倍加速,性能与同尺寸Qwen3模型相当。结果表明,基于家族的参数共享是极有效策略,确立了‘训练一次,部署多次’新范式,为架构效率与高性能平衡提供关键参考。
原文摘要 · Abstract (English)
Large Language Models (LLMs) face significant challenges regarding deployment costs and latency, necessitating adaptive computing strategies. Building upon the AI Flow framework, we introduce Ruyi2 as an evolution of our adaptive model series designed for efficient variable-depth computation. While early-exit architectures offer a viable efficiency-performance balance, the Ruyi model and existing methods often struggle with optimization complexity and compatibility with large-scale distributed training. To bridge this gap, Ruyi2 introduces a stable "Familial Model" based on Megatron-LM. By using 3D parallel training, it achieves a 2-3 times speedup over Ruyi, while performing comparably to same-sized Qwen3 models. These results confirm that family-based parameter sharing is a highly effective strategy, establishing a new "Train Once, Deploy Many" paradigm and providing a key reference for balancing architectural efficiency with high-performance capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。