arXiv:2605.12964cs.CV2026-05被引 2

提出不对称流模型,高效生成高维图像并提升真实感。

Asymmetric Flow Models

论文配图:Asymmetric Flow Models
图 1 · 摘自论文原文
  • 噪声预测限于低秩子空间,数据预测保持全维,实现高效建模。
  • ImageNet上达1.57 FID,显著优于同类像素扩散模型。
  • 可从预训练隐空间模型微调至像素空间,保留高层语义,提升细节真实感。

高维空间中的流模型生成面临挑战,因速度预测需建模高维噪声,即使数据具有强低秩结构。本文提出不对称流建模(AsymFlow),通过将噪声预测限制在低秩子空间,同时保持数据预测为全维,从而实现不对称预测。该方法无需改变网络架构或训练/采样流程,即可解析恢复全维速度。在256×256的ImageNet上,AsymFlow达到领先的1.57 FID,大幅超越先前的DiT/JiT类像素扩散模型。此外,AsymFlow首次实现了将预训练隐空间流模型无缝微调为像素空间模型:对齐像素低秩子空间与隐空间,提供自然初始化,使微调主要聚焦于低级差异修正而非重新学习像素生成。我们展示,基于FLUX.2 klein 9B微调的像素级AsymFlow模型,在HPSv3、DPG-Bench和GenEval上均创下新纪录,且视觉真实感显著提升。

原文摘要 · Abstract (English)

Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has strong low-rank structure. We present Asymmetric Flow Modeling (AsymFlow), a rank-asymmetric velocity parameterization that restricts noise prediction to a low-rank subspace while keeping data prediction full-dimensional. From this asymmetric prediction, AsymFlow analytically recovers the full-dimensional velocity without changing the network architecture or training/sampling procedures. On ImageNet 256$\times$256, AsymFlow achieves a leading 1.57 FID, outperforming prior DiT/JiT-like pixel diffusion models by a large margin. AsymFlow also provides the first-ever route for finetuning pretrained latent flow models into pixel-space models: aligning the low-rank pixel subspace to the latent space gives a seamless initialization that preserves the latent model's high-level semantics and structure, so finetuning mainly improves low-level mismatches rather than relearning pixel generation. We show that the pixel AsymFlow model finetuned from FLUX.2 klein 9B establishes a new state of the art for pixel-space text-to-image generation, beating its latent base on HPSv3, DPG-Bench, and GenEval while qualitatively showing substantially improved visual realism.

流模型图像生成微调真实感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。