轻量级多模态模型,专精企业文档理解与安全检测。
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
- 基于20亿参数的解码器架构,对齐视觉模态提升文档理解能力。
- 在视觉文档基准和LiveXiv测试中表现优异,避免数据污染。
- 开源可商用,训练数据透明,适合企业级部署与安全应用。
我们提出Granite Vision,一种专为企事业场景设计的轻量级多模态大语言模型,具备视觉理解能力。该模型在包含表格、图表、示意图、信息图等内容提取任务的综合性指令跟随数据集上训练,采用仅解码器结构,共20亿参数。模型通过在推理阶段引入稀疏注意力向量的安全分类机制,识别潜在有害输入。尽管模型规模轻量,但在标准视觉文档理解基准及针对新论文持续更新的LiveXiv基准上均表现良好,有效规避测试集污染问题。模型以Apache-2许可证发布,支持研究与商业使用,并提供完整的训练数据与细节透明度。模型权重可在https://huggingface.co/ibm-granite/ 获取。
原文摘要 · Abstract (English)
We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document understanding. Our model is trained on a comprehensive instruction-following dataset, including document-related tasks, such as content extraction from tables, charts, diagrams, sketches, and infographics, as well as general image tasks. The architecture of Granite Vision is centered around visual modality alignment with a decoder-only, 2 billion parameter Granite large language model. Additionally, we introduce a dedicated safety classification approach in test-time that leverages a sparse set of attention vectors to identify potential harmful inputs. Despite its lightweight architecture, Granite Vision achieves strong results in standard benchmarks related to visual document understanding, as well as on the LiveXiv benchmark, which is designed to avoid test set contamination by using a constantly updated corpus of recently published Arxiv papers. We are releasing the model under the Apache-2 license, allowing for both research and commercial use, while offering complete visibility into the training data and other relevant details. See https://huggingface.co/ibm-granite/ for model weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。