TextHawk2用1/16的令牌量实现双语文字识别与图像定位,效率更高。
TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
- 通过令牌压缩与视觉编码器强化,实现低资源下的精细感知。
- 在OCR、图表问答等任务上表现超群,最高准确率达89.6%。
- 适合需要高效多语言图文理解的部署场景,如文档处理。
阅读密集文本并定位图像中物体是大型视觉语言模型完成复杂任务的基础能力。以往的大型视觉语言模型(包括GPT-4o等优秀闭源模型)难以同时在两项任务上表现出色。此外,具备细粒度感知能力的模型每张图像需消耗数千令牌,资源开销大。我们提出TextHawk2,一种双语大型视觉语言模型,具备高效的细粒度感知能力,在通用任务、文字识别(OCR)和图像定位任务上均达到领先水平,且图像令牌数减少16倍。关键改进包括:(1)令牌压缩:基于前代高效架构,将每张图像的令牌数减少16倍,显著降低训练与部署资源需求;(2)视觉编码器增强:通过视觉语言模型协同训练,提升编码器对中文OCR和定位等新任务的泛化能力;(3)数据多样性:保持预训练数据规模约1亿样本的同时,丰富数据来源。我们在多个基准测试中评估TextHawk2,表现持续领先,优于同规模闭源模型,如在OCRBench上达78.4%准确率,ChartQA为81.4%,DocVQA的ANLS达89.6%,RefCOCOg-test的[email protected]为88.1%。
原文摘要 · Abstract (English)
Reading dense text and locating objects within images are fundamental abilities for Large Vision-Language Models (LVLMs) tasked with advanced jobs. Previous LVLMs, including superior proprietary models like GPT-4o, have struggled to excel in both tasks simultaneously. Moreover, previous LVLMs with fine-grained perception cost thousands of tokens per image, making them resource-intensive. We present TextHawk2, a bilingual LVLM featuring efficient fine-grained perception and demonstrating cutting-edge performance across general-purpose, OCR, and grounding tasks with 16 times fewer image tokens. Critical improvements include: (1) Token Compression: Building on the efficient architecture of its predecessor, TextHawk2 significantly reduces the number of tokens per image by 16 times, facilitating training and deployment of the TextHawk series with minimal resources. (2) Visual Encoder Reinforcement: We enhance the visual encoder through LVLM co-training, unlocking its potential for previously unseen tasks like Chinese OCR and grounding. (3) Data Diversity: We maintain a comparable scale of 100 million samples while diversifying the sources of pre-training data. We assess TextHawk2 across multiple benchmarks, where it consistently delivers superior performance and outperforms closed-source models of similar scale, such as achieving 78.4% accuracy on OCRBench, 81.4% accuracy on ChartQA, 89.6% ANLS on DocVQA, and 88.1% [email protected] on RefCOCOg-test.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。