通过多维流匹配加速图像生成,兼顾质量与速度
XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative Modeling

- 采用时空双维度条件控制增强概率路径可学习性
- 实现7.2-8.5倍教师模型加速,FID表现领先
- 适合追求高效高质生成的开发者和研究者
高质量图像生成面临速度与质量的权衡。扩散模型虽视觉效果优异,但需耗时的迭代采样。现有高效方法主要将预训练模型压缩为少步采样器,过程困难且依赖教师模型质量。本文提出XYZFlow,通过多维流匹配重新思考高效生成。不同于单步映射,XYZFlow通过结构化多维条件提升概率路径的可识别性和可学习性。将自回归建模视为隐式流拉直,更丰富的上下文降低轨迹歧义。该思想通过两个正交维度实现:时间缩放,利用完整的去噪历史进行非马尔可夫条件;空间缩放,通过下一快捷预测(Next Shortcut Prediction)依次生成图像块,以前置块的去噪轨迹作为先验。实验表明,XYZFlow在保持竞争力的FID指标下,实现7.2-8.5倍教师模型速度提升,且下一快捷预测在质量-延迟权衡上优于模型缩放或步数减少。
原文摘要 · Abstract (English)
High-fidelity image generation faces a trade-off between speed and quality. Diffusion models produce strong visuals but require costly iterative sampling. Existing efficient methods mainly distill pretrained models into few-step samplers, a challenging process that depends heavily on teacher-model quality. In this paper, we introduce XYZFlow, a framework that rethinks efficient generation through multidimensional scaling of flow matching. Unlike single-step mappings, XYZFlow enhances expressivity by making probability paths more identifiable and learnable through structured multidimensional conditioning. We view autoregressive modeling as implicit flow straightening, where richer context reduces trajectory ambiguity. XYZFlow realizes this idea through two orthogonal dimensions: temporal scaling, which uses non-Markovian conditioning on the full denoising history; and spatial scaling, enabled by Next Shortcut Prediction, which sequentially generates patches using preceding patches' denoising trajectories as priors. Experiments show that XYZFlow achieves state-of-the-art performance, with 7.2-8.5X teacher speedups and competitive FID, while Next Shortcut Prediction delivers superior quality-latency trade-offs over model scaling or step reduction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。