arXiv:2502.17092cs.CVcs.CL2025-02

小模型也能高效处理企业级多模态任务,靠的是设计优化而非海量数据。

Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI

  • 通过注意力稳定性和位置编码改进,减少训练所需数据量。
  • 1B和4B参数模型在文档理解等任务上表现媲美大模型。
  • 适合资源有限但需高效率的企业多模态应用。

我们提出Shakti-VLM系列视觉语言模型,包含10亿和40亿参数规模,旨在解决多模态学习中的数据效率问题。尽管近期模型依赖大量训练数据取得优异性能,但Shakti通过架构创新,在更少的训练令牌下仍能实现竞争力结果。关键技术包括用于注意力稳定的QK归一化、混合归一化策略及增强的位置编码,并采用三阶段训练策略提升学习效率。评估显示,Shakti-VLM-1B与Shakti-VLM-4B在文档理解、视觉推理、OCR提取及通用多模态推理任务中表现卓越。结果表明,高性能可通过模型设计与训练策略实现,而非单纯依赖数据规模,使Shakti成为企业级多模态任务的高效解决方案。

原文摘要 · Abstract (English)

We introduce Shakti VLM, a family of vision-language models in the capacity of 1B and 4B parameters designed to address data efficiency challenges in multimodal learning. While recent VLMs achieve strong performance through extensive training data, Shakti models leverage architectural innovations to attain competitive results with fewer tokens. Key advancements include QK-Normalization for attention stability, hybrid normalization techniques, and enhanced positional encoding. A three-stage training strategy further optimizes learning efficiency. Evaluations show that Shakti-Shakti-VLM-1B and Shakti-VLM-4B excel in document understanding, Visual Reasoning, OCR extraction, and general multimodal reasoning. Our results highlight that high performance can be achieved through model design and training strategy rather than sheer data volume, making Shakti an efficient solution for enterprise-scale multimodal tasks.

视觉语言模型多模态企业AI模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。