arXiv:2605.02641cs.CV2026-05被引 2

Mamoda2.5用专家混合提升多模态生成,提速近百倍且效果顶尖。

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE

论文配图:Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
图 1 · 摘自论文原文
  • 采用细粒度MoE结构,128专家中仅激活3B参数,降低训练成本。
  • 在VBench 2.0和OpenVE-Bench上达顶尖水平,视频编辑质量超开源模型。
  • 4步推理压缩技术使生成速度提升95.9倍,适合实时应用部署。

我们提出Mamoda2.5,一个统一的AR-Diffusion框架,将多模态理解与生成整合于单一架构中。为高效提升生成能力,我们在扩散变压器骨干网络中引入细粒度的专家混合(MoE)设计(128专家,Top-8路由),构建出250亿参数模型,但仅激活30亿参数,显著降低训练成本的同时扩大模型容量。Mamoda2.5在VBench 2.0上达到顶级生成性能,在视频编辑质量上创下新纪录,超越评估的开源模型,媲美当前顶级专有模型,包括Kling O1在OpenVE-Bench上的表现。此外,我们提出联合少步蒸馏与强化学习框架,将30步编辑模型压缩至4步,极大加速推理。相比开源基线,其视频编辑推理速度最快提升95.9倍。在实际应用中,该模型已成功部署于广告场景的内容审核与创意修复任务,内部视频编辑成功率高达98%。

原文摘要 · Abstract (English)

We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the model's generation capability, we equip the Diffusion Transformer backbone with a fine-grained Mixture-of-Experts (MoE) design (128 experts, Top-8 routing), yielding a 25B-parameter model that activates only 3B parameters, significantly reducing training costs while scaling up the model capacity. Mamoda2.5 achieves top-tier generation performance on VBench 2.0 and sets a new record in video editing quality, surpassing evaluated open-source models and matching the performance of current top-tier proprietary models, including the Kling O1 on OpenVE-Bench. Furthermore, we introduce a joint few-step distillation and reinforcement learning framework that compresses the 30-step editing model into a 4-step model and greatly accelerates model inference. Compared to open-source baselines, Mamoda2.5 achieves up to $95.9\times$ faster video editing inference. In real-world applications, Mamoda2.5 has been successfully deployed for content moderation and creative restoration tasks in advertising scenarios, achieving a 98% success rate in internal advertising video editing scenario.

多模态生成MoE视频编辑推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。