arXiv:2511.15059cs.CVcs.CL2025-11被引 5

评测大模型对竖排日文的识别能力,发现现有模型表现不佳。

Evaluating Multimodal Large Language Models on Vertically Written Japanese Text

  • 构建合成日文OCR数据集,含横竖排文本,用于训练与评估。
  • 实测显示模型在竖排日文上表现明显差于横排文本。
  • 用合成数据微调后,不支持竖排的模型性能显著提升。

多模态大语言模型(MLLMs)近年来发展迅速,已应用于视觉文档理解任务,预期可处理跨语言文档图像,包括日文。理解文档图像需准确识别文字内容,而部分日文文档采用竖排书写,因此支持竖排书写至关重要。然而,针对竖排日文的研究仍较有限。本研究评估现有MLLMs对竖排日文文本的阅读能力。首先,通过渲染日文文本生成合成日文OCR数据集,包含横排与竖排文本,用于模型微调与评估;同时构建来自真实文档图像的竖排日文评估数据集。实验表明,现有MLLMs在竖排日文上的表现显著劣于横排日文。此外,使用合成数据集微调后,此前无法处理竖排书写的模型性能得到明显提升。相关数据集与代码已公开:https://github.com/llm-jp/eval_vertical_ja。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have seen rapid advances in recent years and are now being applied to visual document understanding tasks. They are expected to process a wide range of document images across languages, including Japanese. Understanding documents from images requires models to read what are written in them. Since some Japanese documents are written vertically, support for vertical writing is essential. However, research specifically focused on vertically written Japanese text remains limited. In this study, we evaluate the reading capability of existing MLLMs on vertically written Japanese text. First, we generate a synthetic Japanese OCR dataset by rendering Japanese texts into images, and use it for both model fine-tuning and evaluation. This dataset includes Japanese text in both horizontal and vertical writing. We also create an evaluation dataset sourced from the real-world document images containing vertically written Japanese text. Using these datasets, we demonstrate that the existing MLLMs perform worse on vertically written Japanese text than on horizontally written Japanese text. Furthermore, we show that training MLLMs on our synthesized Japanese OCR dataset results in improving the performance of models that previously could not handle vertical writing. The datasets and code are publicly available https://github.com/llm-jp/eval_vertical_ja.

多模态日文识别竖排文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。