120亿参数多模态模型,图像文档理解更强,性能超更大模型。
Pixtral 12B
- 自研视觉编码器,支持原生分辨率图像输入,灵活处理图像令牌数。
- 在多模态基准测试中超越同类开源模型,甚至胜过7倍大的Llama-3.2 90B。
- 兼具强文本能力,适合需要图文理解与生成的开发者与研究者。
我们提出Pixtral-12B,一个120亿参数的多模态语言模型,能够理解自然图像与文档,在多个多模态基准上表现领先,优于许多参数更大的模型。不同于多数开源模型,Pixtral在同等规模下仍具备前沿文本处理能力,未牺牲自然语言性能以换取多模态表现。其采用从头训练的新视觉编码器,可接收原生分辨率和宽高比的图像,用户可灵活控制图像处理的令牌数量。模型支持在128K令牌长上下文窗口内处理任意数量图像。Pixtral-12B显著优于同规模开源模型(如Llama-3.2 11B与Qwen-2-VL 7B),且在参数量仅为7倍的情况下超越大型开源模型(如Llama-3.2 90B)。我们还开源了用于实际场景评估的基准MM-MT-Bench,并提供标准化评估协议的详细分析与代码。Pixtral-12B采用Apache 2.0许可证发布。
原文摘要 · Abstract (English)
We introduce Pixtral-12B, a 12--billion-parameter multimodal language model. Pixtral-12B is trained to understand both natural images and documents, achieving leading performance on various multimodal benchmarks, surpassing a number of larger models. Unlike many open-source models, Pixtral is also a cutting-edge text model for its size, and does not compromise on natural language performance to excel in multimodal tasks. Pixtral uses a new vision encoder trained from scratch, which allows it to ingest images at their natural resolution and aspect ratio. This gives users flexibility on the number of tokens used to process an image. Pixtral is also able to process any number of images in its long context window of 128K tokens. Pixtral 12B substanially outperforms other open models of similar sizes (Llama-3.2 11B \& Qwen-2-VL 7B). It also outperforms much larger open models like Llama-3.2 90B while being 7x smaller. We further contribute an open-source benchmark, MM-MT-Bench, for evaluating vision-language models in practical scenarios, and provide detailed analysis and code for standardized evaluation protocols for multimodal LLMs. Pixtral-12B is released under Apache 2.0 license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。