首个去中心化训练的视频生成模型,实现高质量时序一致生成。
Paris 2.0: A Decentralized Diffusion Model for Video Generation

- 采用去中心化计算训练视频扩散模型,突破传统集群依赖
- 低分辨率文生视频任务下FVD降至279.01,性能提升约2倍
- 适合关注分布式训练与视频生成融合的研究者
我们提出Paris 2.0,首个通过去中心化计算预训练的视频生成模型。其训练方法基于Paris 1.0(arXiv:2510.03434),即首个开源权重的去中心化扩散模型(DDM),已证明图像生成可脱离集中式GPU集群完成。然而,在去中心化训练下实现时序一致的视频生成仍是未解难题,Paris 2.0成功解决该问题。在低分辨率文生视频训练中,与在相同数据和匹配总算力预算下训练的集中式模型相比,Paris 2.0将弗雷谢视频距离(FVD)从561.04降至279.01,提升约2.0倍,并显著提高CLIP文本-视频相似度与美学评分。
原文摘要 · Abstract (English)
We present Paris 2.0, the first video generation model pre-trained through decentralized computation. Its training recipe builds upon Paris 1.0 (arXiv:2510.03434), the first ever open-weight Decentralized Diffusion Model (DDM), which showed that image generation can be trained without a monolithic GPU cluster. However, temporally coherent video generation had remained an open problem under decentralized training, and Paris 2.0 closes it. In low-resolution text-to-video training, against a monolithic model trained on the same data under a matched total compute budget, Paris 2.0 cuts Frechet Video Distance (FVD) from 561.04 to 279.01, a ~2.0x improvement, and lifts CLIP text-video similarity and aesthetic score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。