arXiv:2607.09657cs.CVcs.AI2026-07被引 2

用视觉信息预训练模型,比纯文本更有效提升语言智能。

Scalable Visual Pretraining for Language Intelligence

论文配图:Scalable Visual Pretraining for Language Intelligence
图 1 · 摘自论文原文
  • 直接用图文文档做无监督预训练,不转成纯文本。
  • 在多个数据集上表现优于纯文本预训练,效果更优。
  • 适合需要理解图表、公式等视觉内容的场景。

大型基础模型的快速发展主要依赖大规模文本语料的预训练。然而,许多知识通过视觉形式传递,如图表、排版公式和页面布局,这些信息无法仅靠文本完整表达。当前的预训练方法通常将图文丰富的文档(如网页、论文)转化为纯文本,丢弃了关键视觉线索。本文挑战了‘语言模型必须基于纯文本’的默认假设,证明视觉预训练是可扩展的基础模型智能学习方式。我们系统研究了无需文本提取的无监督视觉预训练范式,直接利用视觉文档进行训练。在多种骨干网络和基准测试中,相同语料库下的视觉预训练持续优于纯文本预训练,为可扩展的语言智能提供了高效路径。

原文摘要 · Abstract (English)

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.

视觉预训练基础模型图文理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。