arXiv:2603.15132cs.CV2026-03被引 1

通过语义路标解决像素空间生成中的轨迹冲突问题

WiT: Waypoint Diffusion Transformers for Alleviating Trajectory Conflict in Pixel-Space Image Generation

  • 引入预训练视觉特征投影的语义路标,重构像素生成路径
  • 在不同目标间保持精准传输方向,实现全像素空间生成
  • 适合追求高质量图像生成且关注生成可控性的研究者

尽管近期的流匹配模型通过直接在像素空间操作避免了潜在自编码器的重建瓶颈,但原始像素流形缺乏显式语义结构,导致共享有限容量向量场难以区分特定目标的传输方向。当不同传输需求在此处局部纠缠时,学习信号会相互干扰,阻碍优化,这种现象称为轨迹冲突。为缓解轨迹冲突并保留像素空间直接生成能力,我们提出路径扩散变压器(WiT),通过从预训练视觉表示中投影出紧凑的语义路标,构建语义组织的中间表示,以补充直接像素预测。在常微分方程积分过程中,轻量级路标预测器从当前噪声状态动态推断语义路标,并通过仅像素自适应归一化机制持续调节主像素生成器。这种空间变化的语义调制重新组织了有限容量学习问题,在保持全像素空间生成的同时,帮助模型保留目标特异性传输方向。ImageNet 2012实验表明,不同规模模型上均优于强基线,其中WiT-H/16取得FID 1.79。

原文摘要 · Abstract (English)

While recent Flow Matching models avoid the reconstruction bottlenecks of latent autoencoders by operating directly in pixel space, the raw pixel manifold provides little explicit semantic organization, making target-specific transport directions difficult for a shared finite-capacity vector field to distinguish. When different transport requirements are locally entangled in this way, their learning signals can interfere and hinder optimization, a phenomenon we refer to as trajectory conflict. To alleviate trajectory conflict while preserving direct generation in pixel space, we propose Waypoint Diffusion Transformers (WiT), which introduces explicit semantic routing into pixel-space generation. WiT structures pixel-space prediction through compact semantic waypoints projected from pre-trained vision representations, providing a semantically organized intermediate representation that complements direct pixel prediction. During ODE integration, a lightweight Waypoints Predictor dynamically infers the semantic waypoint from the current noisy state, and the predicted waypoint continuously conditions the primary Pixel Space Generator through our Just-Pixel AdaLN mechanism. This spatially varying semantic modulation reorganizes the finite-capacity learning problem and helps the model preserve target-specific transport directions while retaining generation entirely in pixel space. Experiments on ImageNet 2012 demonstrate consistent improvements over strong pixel-space baselines across model scales, with WiT-H/16 achieving an FID of 1.79. Code will be released at https://hainuo-wang.github.io/WiT.

扩散模型图像生成语义路由像素空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。