arXiv:2603.25249cs.CV2026-03

让图像生成的潜在表示更懂语义,用少样本实现高质量生成。

Semantic-Aware Prefix Learning for Token-Efficient Image Generation

  • 在查询式1D分块框架中注入类别级语义条件。
  • 通过尾部令牌丢弃策略强制语义信息主导表示学习。
  • 适合追求低资源高效图像生成的研究者与开发者。

视觉分块器在潜在空间图像生成中起核心作用,连接高维图像与可处理的生成建模。然而,大多数现有分块器仍以重建为主导目标,导致潜在表示与高层语义关联较弱。近期方法虽提升语义对齐,但通常将语义信号视为辅助正则化而非必要成分。本文提出SMAP——一种语义感知前缀分块器,将类别级语义条件引入基于查询的1D分块框架。为使语义在训练中不可或缺,SMAP采用尾部令牌丢弃策略,在逐渐缩减的令牌预算下,迫使语义条件和早期潜在前缀承担更大责任。为验证所得潜在空间不仅适用于重建,还利于生成,我们进一步设计了CARD——一种混合因果自回归-扩散生成器。在ImageNet上的大量实验表明,SMAP在离散与连续分块设置下均持续提升重建质量,且其语义驱动的潜在空间在紧凑令牌预算下表现出强大的下游生成性能。

原文摘要 · Abstract (English)

Visual tokenizers play a central role in latent image generation by bridging high-dimensional images and tractable generative modeling. However, most existing tokenizers are still trained with reconstruction-dominated objectives, which often yield latent representations that are only weakly grounded in high-level semantics. Recent approaches improve semantic alignment, but typically treat semantic signals as auxiliary regularization rather than making them functionally necessary for representation learning. We propose SMAP, a SeMantic-Aware Prefix tokenizer that injects class-level semantic conditions into a query-based 1D tokenization framework. To make semantics indispensable during training, SMAP introduces a tail token dropping strategy, which forces semantic conditions and early latent prefixes to bear increasing responsibility under progressively reduced token budgets. To verify that the resulting latent space is useful for generation rather than reconstruction alone, we further introduce CARD, a hybrid Causal AutoRegressive--Diffusion generator. Extensive experiments on ImageNet show that SMAP consistently improves reconstruction quality across discrete and continuous tokenization settings, and that its semantically grounded latent space yields strong downstream generation performance under compact token budgets.

图像生成语义感知低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。