arXiv:2412.15523cs.CVcs.AI2024-12AAAI被引 11

用人类语言指令提升图像文字识别准确率

InstructOCR: Instruction Boosting Scene Text Spotting

  • 引入人类指令增强图文理解,结合文本与图像编码器
  • 在TextVQA和ST-VQA上分别提升2.6%和2.1%性能
  • 适用于文字识别与视觉问答,适合多模态研究者

在场景文字定位任务中,以往的OCR方法主要依赖图像编码器和预训练文本信息,但常忽视人类语言指令的优势。为此,我们提出InstructOCR,一种基于指令的场景文字识别模型,利用人类语言指令增强对图像中文本的理解。该框架在训练和推理中同时使用文本与图像编码器,并基于文本属性精心设计指令,使模型能更准确、灵活地解析文本。大量实验表明,该模型在主流基准上达到当前最优性能。此外,该框架可无缝应用于场景文字视觉问答(VQA)任务。通过在预训练阶段引入指令策略,下游VQA任务性能显著提升,在TextVQA数据集上提高2.6%,在ST-VQA数据集上提高2.1%。这些结果揭示了将人类语言指令融入OCR任务的潜力。

原文摘要 · Abstract (English)

In the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene text spotting model that leverages human language instructions to enhance the understanding of text within images. Our framework employs both text and image encoders during training and inference, along with instructions meticulously designed based on text attributes. This approach enables the model to interpret text more accurately and flexibly. Extensive experiments demonstrate the effectiveness of our model and we achieve state-of-the-art results on widely used benchmarks. Furthermore, the proposed framework can be seamlessly applied to scene text VQA tasks. By leveraging instruction strategies during pre-training, the performance on downstream VQA tasks can be significantly improved, with a 2.6% increase on the TextVQA dataset and a 2.1% increase on the ST-VQA dataset. These experimental results provide insights into the benefits of incorporating human language instructions for OCR-related tasks.

文字识别指令学习多模态VQA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。