arXiv:2410.13611cs.CVcs.AI2024-10被引 10

小模型H2OVL-Mississippi在文档视觉理解上表现卓越,适合手机端使用。

H2OVL-Mississippi Vision Language Models Technical Report

  • 用3700万图文对训练,仅需240小时算力,8张H100跑出小模型
  • 0.8B模型在OCRBench文本识别任务上超越更大模型,2B版通用性能强
  • 开源免费,适合隐私敏感、设备端部署的文档智能应用

小型视觉语言模型(VLM)因可在消费级硬件上高效运行,日益成为隐私保护型、本地化应用的关键。本文提出H2OVL-Mississippi,基于3700万图像-文本对,在8张H100 GPU上使用240小时算力训练。其中,0.8亿参数的H2OVL-Mississippi-0.8B专注于文本识别,在OCRBench的文本识别任务中达到领先水平,优于许多更大模型;另发布20亿参数的H2OVL-Mississippi-2B,适用于通用场景,在多个学术基准上表现优异。两模型均基于此前的H2O-Danube语言模型扩展至视觉领域。模型以Apache 2.0许可证开源,推动文档AI与视觉大模型的普及。

原文摘要 · Abstract (English)

Smaller vision-language models (VLMs) are becoming increasingly important for privacy-focused, on-device applications due to their ability to run efficiently on consumer hardware for processing enterprise commercial documents and images. These models require strong language understanding and visual capabilities to enhance human-machine interaction. To address this need, we present H2OVL-Mississippi, a pair of small VLMs trained on 37 million image-text pairs using 240 hours of compute on 8 x H100 GPUs. H2OVL-Mississippi-0.8B is a tiny model with 0.8 billion parameters that specializes in text recognition, achieving state of the art performance on the Text Recognition portion of OCRBench and surpassing much larger models in this area. Additionally, we are releasing H2OVL-Mississippi-2B, a 2 billion parameter model for general use cases, exhibiting highly competitive metrics across various academic benchmarks. Both models build upon our prior work with H2O-Danube language models, extending their capabilities into the visual domain. We release them under the Apache 2.0 license, making VLMs accessible to everyone, democratizing document AI and visual LLMs.

视觉语言模型文档理解小模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。