arXiv:2412.03555cs.CV2024-12被引 244

PaliGemma 2升级多尺寸视觉语言模型,支持更广任务迁移。

PaliGemma 2: A Family of Versatile VLMs for Transfer

  • 融合SigLIP-So400m与Gemma 2系列,覆盖2B至27B参数规模。
  • 在224/448/896像素三分辨率训练,提升迁移性能。
  • 新扩展OCR、医学报告生成等任务,表现达当前最优。

PaliGemma 2 是基于 Gema 2 系列语言模型的开源视觉-语言模型(VLM)升级版。我们结合 SigLIP-So400m 视觉编码器与从 2B 到 27B 参数的全部 Gemma 2 模型,分三个分辨率(224px、448px、896px)进行多阶段训练,以增强其通过微调实现广泛迁移的能力。该模型家族涵盖不同规模与分辨率,可用于研究影响迁移性能的因素(如学习率),并分析任务类型、模型大小与分辨率之间的相互作用。此外,其迁移任务范围超越 PaliGemma,新增表格结构识别、分子结构识别、乐谱识别、长细粒度图像描述及放射科报告生成等,均取得当前最优结果。

原文摘要 · Abstract (English)

PaliGemma 2 is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models. We combine the SigLIP-So400m vision encoder that was also used by PaliGemma with the whole range of Gemma 2 models, from the 2B one all the way up to the 27B model. We train these models at three resolutions (224px, 448px, and 896px) in multiple stages to equip them with broad knowledge for transfer via fine-tuning. The resulting family of base models covering different model sizes and resolutions allows us to investigate factors impacting transfer performance (such as learning rate) and to analyze the interplay between the type of task, model size, and resolution. We further increase the number and breadth of transfer tasks beyond the scope of PaliGemma including different OCR-related tasks such as table structure recognition, molecular structure recognition, music score recognition, as well as long fine-grained captioning and radiography report generation, on which PaliGemma 2 obtains state-of-the-art results.

视觉语言模型迁移学习多模态OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。