证明去中心化自回归生成与集中式训练理论等价,提升可扩展性。
Decentralized Autoregressive Generation
- 用离散流匹配框架重构自回归生成,实现模型自然分解为独立专家
- 实验证明去中心化训练在多模态基准上性能媲美集中式架构
- 为去中心化生成提供理论支撑,适合大规模模型开发者参考
近年来,自回归生成的去中心化架构因其对扩展瓶颈的缓解而受到广泛关注。然而,尽管实证效果良好,该范式目前仍缺乏严格的理论基础。本文首次形式化建立了去中心化与集中式训练之间的理论等价性。为此,我们针对自回归生成适配了离散流匹配(Discrete Flow Matching)框架,利用其内在性质证明全局模型可自然分解为独立专家。最后,我们在多个异构多模态基准上进行了广泛实验,实证验证了去中心化训练在性能上与标准集中式架构保持竞争力。
原文摘要 · Abstract (English)
The decentralization of autoregressive generation has attracted considerable attention in recent years as a solution to scaling bottlenecks. However, despite promising empirical results, this paradigm currently lacks rigorous theoretical justification. In this work, we formally establish the theoretical equivalence between decentralized and centralized training. To achieve this, we adapt the Discrete Flow Matching framework for autoregressive generation, leveraging its inherent properties to demonstrate that global models naturally decompose into independent experts. Finally, we conduct extensive experiments across diverse multimodal benchmarks, empirically validating that decentralized training maintains competitive parity with standard centralized architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。