arXiv:2509.23122cs.CV2025-09被引 6

提出新型耦合生成框架,兼顾图像质量与计算效率。

Stochastic Interpolants via Conditional Dependent Coupling

  • 基于条件依赖耦合构建多阶段生成路径
  • 在多个分辨率下实现高保真与高效生成
  • 适合需要端到端训练的图像生成任务

现有图像生成模型面临计算成本与生成质量之间的权衡。依赖预训练变分自编码器(VAE)的模型存在信息丢失、细节有限且无法支持端到端训练的问题;而直接在像素空间操作的模型则计算开销巨大。尽管级联模型可缓解计算压力,但各阶段分离导致难以实现端到端优化、知识共享受限,并常造成各阶段分布学习不准。为此,本文提出基于条件依赖耦合的统一多阶段生成框架,将生成过程分解为多阶段插值轨迹,确保分布学习准确的同时支持端到端优化。整个过程被建模为单一扩散变换器(Diffusion Transformer),消除了模块间的割裂,实现知识共享。大量实验表明,该方法在多种分辨率下均实现了高质量与高效率的平衡。

原文摘要 · Abstract (English)

Existing image generation models face critical challenges regarding the trade-off between computation and fidelity. Specifically, models relying on a pretrained Variational Autoencoder (VAE) suffer from information loss, limited detail, and the inability to support end-to-end training. In contrast, models operating directly in the pixel space incur prohibitive computational cost. Although cascade models can mitigate computational cost, stage-wise separation prevents effective end-to-end optimization, hampers knowledge sharing, and often results in inaccurate distribution learning within each stage. To address these challenges, we introduce a unified multistage generative framework based on our proposed Conditional Dependent Coupling strategy. It decomposes the generative process into interpolant trajectories at multiple stages, ensuring accurate distribution learning while enabling end-to-end optimization. Importantly, the entire process is modeled as a single unified Diffusion Transformer, eliminating the need for disjoint modules and also enabling knowledge sharing. Extensive experiments demonstrate that our method achieves both high fidelity and efficiency across multiple resolutions.

图像生成扩散模型端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。