arXiv:2609.09143cs.CVcs.CL2026-09

通过联合建模评估图像分词器,揭示其与文本的协同学习机制。

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

论文配图:Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
图 1 · 摘自论文原文
  • 构建纯自回归测试环境,追踪多任务预训练中的损失变化。
  • 图像到文本损失对生成与理解性能均有稳定预测能力。
  • 分词器设计影响文本建模,重建效果不等于下游表现提升。

图像分词器定义了统一多模态模型的‘视觉语言’,但通常仅通过孤立指标或生成/理解单一任务评估。这些方法无法充分反映视觉标记在与文本联合建模时的行为。本文构建一个受控的纯自回归测试平台,在文本、图像、文生图(T2I)和图生文(I2T)预测的持续预训练过程中追踪任务特定验证损失。分析损失随规模的变化及其与下游性能的关系,进而研究多模态可学习性及分词器设计。发现:(1)损失需按任务分析,因不同任务呈现不同缩放行为,且分词器排序不同;(2)损失-性能关系依赖于目标标记空间:固定分词器下,T2I与I2T损失与生成质量相关;跨分词器时,T2I损失关系随图像标记空间变化,而基于共享文本词汇的I2T损失则更一致,且在微调后与生成和视觉理解性能均相关。以损失为视角,发现:(3)更好重建未必带来更低任务损失或更强下游性能;(4)分词器选择会影响联合优化下的文本建模。作为案例,重新考察判别器、语义监督与词汇量三个设计维度对联合建模与下游性能的影响。整体提供了一种互补视角,揭示图像分词器作为视觉语言与文本在联合训练中的相互作用。

原文摘要 · Abstract (English)

Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. We examine how these losses scale and relate to downstream performance, then use them to study multimodal learnability---how well image and text tokens are jointly modeled---and tokenizer design. We find that (1) losses should be analyzed by task, since they exhibit distinct scaling behavior and rank tokenizers differently. (2) The loss--performance relationship depends on the predicted token space: for a fixed tokenizer, T2I and I2T losses correlate with generation quality, but across tokenizers, the T2I loss--performance relationship shifts with the image-token space, whereas I2T loss, computed over a shared text vocabulary, provides a more consistent signal. I2T loss also correlates with both generation and visual understanding performance after supervised finetuning. Using losses as a lens, we show that (3) better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that (4) image tokenizer choice can affect text modeling under joint optimization. As case studies, we revisit three tokenizer design axes---the discriminator, semantic supervision, and vocabulary size---to examine their effects on joint modeling and downstream performance. Together, our testbed offers a complementary perspective on image tokenizers as visual languages, highlighting their interplay with text in joint multimodal training.

图像分词器多模态建模联合训练生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。