arXiv:2509.26625cs.LGcs.AI2025-09被引 24

语言模型通过文本训练也能获得视觉知识,且可分出感知与推理两类先验。

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

  • 从语言数据中分离出视觉感知和推理两类先验,机制不同
  • 推理先验随推理类数据(如代码、数学)增加而持续提升
  • 适合想构建视觉能力的多模态大模型的研究者参考

大型语言模型(LLMs)虽仅用文本训练,却意外具备丰富的视觉先验。这些先验使模型在少量多模态数据下即可解锁视觉能力,甚至无需见过图像即可完成视觉任务。通过系统分析,我们发现视觉先验由可分离的感知与推理先验组成,具有不同增长趋势和来源。模型的潜在视觉推理能力主要来自以推理为中心的数据(如代码、数学、学术文献)预训练,并呈渐进式增长;这类推理先验可迁移且通用。相比之下,感知先验更广泛地源于广义语料库,且对视觉编码器和视觉指令调优数据更敏感。同时,描述视觉世界的文本至关重要,但其影响迅速饱和。基于此,我们提出一种以数据为核心的视觉感知型语言模型预训练方案,并在1万亿令牌规模下验证。研究基于超过100项受控实验,耗时50万GPU小时,覆盖从语言模型预训练到视觉对齐及监督微调的全流程,涵盖五种模型规模、多种数据类别与组合、多重适配设置。此外,我们提出并检验多个假设,引入多层级存在基准(MLE-Bench)。本工作为从语言预训练中主动培养视觉先验提供了新路径,推动下一代多模态大模型发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs), despite being trained on text alone, surprisingly develop rich visual priors. These priors allow latent visual capabilities to be unlocked for vision tasks with a relatively small amount of multimodal data, and in some cases, to perform visual tasks without ever having seen an image. Through systematic analysis, we reveal that visual priors-the implicit, emergent knowledge about the visual world acquired during language pre-training-are composed of separable perception and reasoning priors with unique scaling trends and origins. We show that an LLM's latent visual reasoning ability is predominantly developed by pre-training on reasoning-centric data (e.g., code, math, academia) and scales progressively. This reasoning prior acquired from language pre-training is transferable and universally applicable to visual reasoning. In contrast, a perception prior emerges more diffusely from broad corpora, and perception ability is more sensitive to the vision encoder and visual instruction tuning data. In parallel, text describing the visual world proves crucial, though its performance impact saturates rapidly. Leveraging these insights, we propose a data-centric recipe for pre-training vision-aware LLMs and verify it in 1T token scale pre-training. Our findings are grounded in over 100 controlled experiments consuming 500,000 GPU-hours, spanning the full MLLM construction pipeline-from LLM pre-training to visual alignment and supervised multimodal fine-tuning-across five model scales, a wide range of data categories and mixtures, and multiple adaptation setups. Along with our main findings, we propose and investigate several hypotheses, and introduce the Multi-Level Existence Bench (MLE-Bench). Together, this work provides a new way of deliberately cultivating visual priors from language pre-training, paving the way for the next generation of multimodal LLMs.

视觉先验多模态大模型推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。