arXiv:2605.14333cs.CV2026-05

针对离散化图像生成中文字和人脸保真度低的问题,提出新分词方法提升细节还原能力。

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation

论文配图:InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
图 1 · 摘自论文原文
  • 设计局部感知损失,让分词器专注保留文字与人脸的精细结构。
  • 在16×下采样、16k代码本下,文字和人脸重建质量显著优于现有方法。
  • 适用于需要高保真文本和人脸的自回归图像生成场景。

文字和人脸是视觉生成中最显著且实用的视觉模式,但在基于离散分词的自回归生成器中仍面临挑战。核心瓶颈在于分词器:激进的下采样与量化常丢失保持可读字符和独特面部特征所需的细粒度结构。我们发现,标准离散分词目标与文字可读性和面部保真度对齐较弱,因其通常优化通用重建并均匀压缩多样化内容。为此,我们提出InsightTok,一种简单但有效的离散视觉分词框架,通过局部化、内容感知的感知损失增强文字和人脸保真度。采用紧凑的16k代码本和16x下采样率,InsightTok在不牺牲整体重建质量的前提下,显著优于先前分词器的文字与人脸重建表现。这些优势在InsightAR自回归图像生成中持续体现,生成图像具备更清晰的文字与更忠实的面部细节。结果表明,分词器训练中引入特定监督对推进离散图像生成具有潜力。

原文摘要 · Abstract (English)

Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central bottleneck is the tokenizer: aggressive downsampling and quantization often discard the fine-grained structures needed to preserve readable glyphs and distinctive facial features. We attribute this gap to standard discrete-tokenizer objectives being weakly aligned with text legibility and facial fidelity, as these objectives typically optimize generic reconstruction while compressing diverse content uniformly. To address this, we propose InsightTok, a simple yet effective discrete visual tokenization framework that enhances text and face fidelity through localized, content-aware perceptual losses. With a compact 16k codebook and a 16x downsampling rate, InsightTok significantly outperforms prior tokenizers in text and face reconstruction without compromising general reconstruction quality. These gains consistently transfer to autoregressive image generation in InsightAR, producing images with clearer text and more faithful facial details. Overall, our results highlight the potential of specialized supervision in tokenizer training for advancing discrete image generation.

图像生成分词器人脸保真文字识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。