arXiv:2502.09093cs.CV2025-02被引 1

让图像和文字在模型中对齐,提升多模态理解能力。

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs

  • 用视觉动态嵌入引导预训练,让图像特征参与自回归建模。
  • 13个基准测试中表现优于现有方法,显著提升对齐精度。
  • 无需修改模型结构,适合各类主流多模态大模型使用。

尽管多模态大模型在感知任务上表现良好,但其多模态对齐精度不足,制约了性能提升。为解决该问题,本文提出视觉动态嵌入引导的预训练(VDEP),一种融合自回归训练的混合范式。通过利用视觉编码器后接MLP生成的动态嵌入,监督图像隐藏状态,并将图像令牌融入自回归训练过程。现有模型多聚焦于从文本输入中恢复信息,常忽视对图像数据的有效处理。本文创新性地将多模态对齐重定义为从输入数据中恢复信息的过程,尤其强调对细节视觉特征的重建。该方法可无缝集成至标准模型中,无需架构改动。在13个基准测试上的实验表明,VDEP优于基线方法,显著超越现有技术。

原文摘要 · Abstract (English)

While MLLMs perform well on perceptual tasks, they lack precise multimodal alignment, limiting performance. To address this challenge, we propose Vision Dynamic Embedding-Guided Pretraining (VDEP), a hybrid autoregressive training paradigm for MLLMs. Utilizing dynamic embeddings from the MLP following the visual encoder, this approach supervises image hidden states and integrates image tokens into autoregressive training. Existing MLLMs primarily focused on recovering information from textual inputs, often neglecting the effective processing of image data. In contrast, the key improvement of this work is the reinterpretation of multimodal alignment as a process of recovering information from input data, with particular emphasis on reconstructing detailed visual features.The proposed method seamlessly integrates into standard models without architectural changes. Experiments on 13 benchmarks show VDEP outperforms baselines, surpassing existing methods.

多模态对齐自回归预训练视觉嵌入MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。