arXiv:2605.16409cs.CVcs.CL2026-05被引 3

让大模型更准地读多语种文字,不依赖外部工具

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models

论文配图:Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
图 1 · 摘自论文原文
  • 用合成文本和多语言数据训练模型,提升图文定位能力
  • 文字识别完整率从71.3升至84.6,幻觉率降至5.5%
  • 适合需要精准读图中文字的多模态应用

光学字符识别(OCR)与多语言场景文本理解对多模态大模型仍是挑战,尤其在小字、模糊、遮挡、手写和复杂排版的现实图像中。我们提出一种无需外部OCR引擎、推理时无需文本框或提取文本的多语言后训练框架。该框架结合约500万条新增多语言训练样本、可控合成OCR生成、图像内文本翻译、基于LoRA的监督微调(SFT)及轻量级OCR导向思维链提示。在独立的真实世界多语言OCR基准测试中,该方法将文字识别完整率从71.3提升至84.6,幻觉率从18.3%降至5.5%,翻译BLEU-1从52.3提高到80.2,且在模糊与旋转条件下仍显著降低幻觉。公开基准评估显示其在强依赖OCR任务上取得提升,同时保持整体多模态能力;消融实验表明SFT为性能提升主因,提示仅提供辅助增益。结果证明,以数据为中心的OCR感知后训练是提升通用多模态大模型多语言图文对齐能力的实用且可扩展方案。

原文摘要 · Abstract (English)

Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images containing small or degraded text, cluttered layouts, occlusion, handwriting, and complex typography. We present an OCR-aware multilingual post-training framework that improves visual-text grounding in a general-purpose MLLM without requiring an external OCR engine, OCR-extracted text, or text bounding boxes at inference time. The framework combines large-scale multilingual OCR supervision, approximately 5M additional multilingual training samples, controlled synthetic OCR generation and in-image text translation, LoRA-based supervised fine-tuning (SFT), and lightweight OCR-oriented Chain-of-Thought prompting. On a held-out real-world multilingual OCR benchmark, OCR-SFT improves OCR completeness from 71.3 to 84.6, reduces hallucination rate from 18.3\% to 5.5\%, and improves translation BLEU-1 from 52.3 to 80.2, with substantial hallucination reductions under blur and rotation. Evaluation on public benchmarks further shows gains on OCR-intensive tasks while largely preserving broader multimodal capabilities; ablations show that SFT provides the primary improvement, with prompting offering smaller complementary gains. These results demonstrate that data-centric OCR-aware post-training provides a practical and scalable approach to improving multilingual visual-text grounding in general-purpose MLLMs.

多语言识别视觉理解大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。