arXiv:2510.02543cs.CV2025-10被引 1

用OCR增强视觉语言模型,提升韩英双语VQA表现

Exploring OCR-augmented Generation for Bilingual VQA

  • 训练1亿样本的KLOCR双语OCR模型,增强VLM文字理解能力
  • 韩语VQA任务上,使用OCR文本使模型准确率显著提升
  • 适合多语言视觉问答、OCR与生成融合的研究者参考

本文研究视觉语言模型(VLMs)在韩英双语场景下的OCR增强生成方法。为推动该领域发展,我们基于1亿条实例训练并发布了KLOCR——一个强大的双语OCR基线模型,可为VLMs提供文本理解能力。同时,我们构建了针对韩语VQA的基准KOCRBench,分析不同提示策略的影响。大量实验表明,利用OCR提取的文本能显著提升开源及商用模型在多语言场景下的表现。本工作为双语视觉问答中的OCR增强生成提供了新见解。模型、代码与数据已公开于https://github.com/JHLee0513/KLOCR。

原文摘要 · Abstract (English)

We investigate OCR-augmented generation with Vision Language Models (VLMs), exploring tasks in Korean and English toward multilingualism. To support research in this domain, we train and release KLOCR, a strong bilingual OCR baseline trained on 100M instances to augment VLMs with OCR ability. To complement existing VQA benchmarks, we curate KOCRBench for Korean VQA, and analyze different prompting methods. Extensive experiments show that OCR-extracted text significantly boosts performance across open source and commercial models. Our work offers new insights into OCR-augmented generation for bilingual VQA. Model, code, and data are available at https://github.com/JHLee0513/KLOCR.

多语言VQAOCR增强视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。