通过联合去噪高低分辨率潜空间,实现高分辨率图像的细节与布局协同生成。
JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation

- 双流并行去噪:低分辨率控制布局,高分辨率控制细节,实时交互融合信息。
- 在相同采样步数下,生成图像细节更丰富,视觉质量优于现有方法。
- 适合需要精细控制构图与纹理的艺术创作场景,如数字绘画、游戏美术设计。
尽管文本到图像生成模型已取得显著成果,但在生成密集细节的高分辨率(HR)图像方面仍存在挑战。当前主流方法采用从低分辨率(LR)到高分辨率的渐进式生成策略:先生成低分辨率图像,再以该图像为引导生成上采样的高分辨率版本。本文提出联合潜空间轨迹(JoLT),在每一步采样中同时对低分辨率和高分辨率潜变量进行去噪。其中,低分辨率潜变量负责控制整体布局,高分辨率潜变量负责生成细节。两个分支通过交叉连接机制实时共享信息,实现双向优化。我们在多个基准上进行了充分验证,结果表明,相比现有基线方法,该方法生成的图像不仅细节更丰富、结构更合理,且视觉效果更佳,为艺术创作提供了新可能。
原文摘要 · Abstract (English)
Although text-to-image generative models produce impressive results, they struggle to generate densely detailed, high-resolution (HR) images. Current literature addresses this issue with a low-to-high-resolution approach. First, a low-resolution (LR) image is generated. Then, an upsampled version is generated using the LR image as an additional cue. In this paper, we present Joint Latent Trajectories (JoLT). To generate an image, JoLT uses two streams that jointly denoise LR and HR latent images at each sampling step. The LR latent controls the overall layout, while the HR latent controls the details. We interconnect both branches to jointly integrate their information. We extensively validate our method, demonstrating its advantages over competing baselines. The resulting images are not only richly detailed but also visually pleasing, opening new avenues for artistic creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。