arXiv:2512.23512cs.CLcs.AI2025-12被引 5

生成任务能提升视觉理解,但仅当在语义层面进行时有效。

UniHetero: Could Generation Enhance Understanding for Vision-Language-Model at Large Data Scale?

  • 在语义层面进行视觉自回归,而非像素级生成
  • 200万+样本下,生成任务提升理解性能并优化数据利用率
  • 适用于多模态扩展,可提升细节捕捉能力

视觉语言大模型正向统一视觉理解与生成任务发展。然而,在大规模预训练(>200M样本)下,生成是否能增强理解仍不明确。本文通过简洁模型UniHetero分析发现:(1)生成仅在语义层面有助于理解,即模型在大型语言模型中自回归高层视觉表示时有效;一旦像素级目标(如扩散损失)直接干扰语言模型,理解性能反而下降。(2)统一生成-理解任务展现出更优的数据缩放趋势和更高数据利用率,表明从视觉模态直接学习视觉知识比依赖图文对齐更高效。(3)对输入嵌入进行自回归能更好捕捉视觉细节,误差累积更少且模态无关,可推广至所有模态。所学语义表示包含物体、位置、形状、颜色等信息,并支持像素级图像生成。

原文摘要 · Abstract (English)

Vision-language large models are moving toward the unification of visual understanding and visual generation tasks. However, whether generation can enhance understanding is still under-explored on large data scale. In this work, we analysis the unified structure with a concise model, UniHetero, under large-scale pretraining (>200M samples). Our key observations are: (1) Generation can improve understanding, but Only if you generate Semantics, Not Pixels. A common assumption in unified vision-language models is that adding generation will naturally strengthen understanding. However, this is not always true at scale. At 200M+ pretraining samples, generation helps understanding only when it operates at the semantic level, i.e. when the model learns to autoregress high-level visual representations inside the LLM. Once pixel-level objectives (e.g., diffusion losses) directly interfere with the LLM, understanding performance often degrades. (2) Generation reveals a superior Data Scaling trend and higher Data Utilization. Unified generation-understanding demonstrates a superior scaling trend compared to understanding alone, revealing a more effective way to learn vision-only knowledge directive from vision modality rather than captioning to text. (3) Autoregression on Input Embedding is effective to capture visual details. Compared to the commonly-used vision encoder, make visual autoregression on input embedding shows less cumulative error and is modality independent, which can be extend to all modalities. The learned semantic representations capture visual information such as objects, locations, shapes, and colors; further enable pixel-level image generation.

视觉语言模型生成增强理解大规模预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。