arXiv:2512.04032cs.CLcs.AI2025-12被引 1

24亿参数小模型实现多语言视觉问答新纪录

jina-vlm: Small Multilingual Vision Language Model

  • 用图像拼贴+注意力池化,高效处理任意分辨率图像
  • 在20个语言上达成开源20亿级模型最高多语言问答准确率
  • 首次系统分析训练数据各类别作用,指导数据构建

我们提出 jina-vlm,一个仅24亿参数的高效多语言视觉语言模型,在开源20亿级模型中实现了领先的多语言视觉问答性能。该模型结合SigLIP2视觉编码器与Qwen3语言解码器,采用图像拼贴和注意力池化技术,实现对任意分辨率图像的高效处理。为分析不同训练数据类别贡献,我们进行了留一法数据混合消融实验,系统性地移除任务、领域、模态和语言类别,诊断各类数据的必要性与冗余性,并考察任务收益在不同领域的迁移能力。模型权重与代码已公开发布于 https://huggingface.co/jinaai/jina-vlm。

原文摘要 · Abstract (English)

We present jina-vlm, a token-efficient 2.4B parameter vision-language model that achieves state-of-the-art multilingual VQA performance among open 2B-scale VLMs. The model couples a SigLIP2 vision encoder with a Qwen3 language decoder and makes use of image tiling and attention-pooling for token-efficient processing of arbitrary-resolution images. To understand the contribution of different training data categories, we conduct a leave-one-out data mixture ablation study-systematically removing task, domain, modality, and language categories-to diagnose which data types are necessary versus redundant and whether task benefits transfer across domains. Model weights and code are publicly released at https://huggingface.co/jinaai/jina-vlm.

多语言视觉语言模型小模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。