arXiv:2502.14778cs.CLcs.AI2025-02ACL被引 1

用日文PDF自动构建多模态训练数据,提升日语大模型表现

Harnessing PDF Data for Improving Japanese Large Multimodal Models

  • 自动化提取日文PDF中的图文对,免去人工标注
  • 在Heron-Bench上性能提升2.1%至13.8%
  • 适合需要提升日语多模态理解能力的研究者

大型多模态模型(LMMs)在英语任务中表现优异,但在日语场景下仍受限于高质量训练数据的缺乏。当前日语LMM多依赖英文数据翻译,难以捕捉日本本土文化知识。本文探索日文PDF数据作为训练资源的潜力,提出一种全自动流水线,利用预训练模型通过版面分析、OCR与视觉-语言配对技术,自动提取图文对,无需人工标注。同时,基于提取结果构建指令数据以丰富训练集。在日语LMM基准测试中,使用该数据训练的模型在Heron-Bench上实现2.1%至13.8%的性能提升。进一步分析表明,该数据对模型规模与语言模型选择均有显著影响,证实其作为日语多模态训练资源的重要价值。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have demonstrated strong performance in English, but their effectiveness in Japanese remains limited due to the lack of high-quality training data. Current Japanese LMMs often rely on translated English datasets, restricting their ability to capture Japan-specific cultural knowledge. To address this, we explore the potential of Japanese PDF data as a training resource, an area that remains largely underutilized. We introduce a fully automated pipeline that leverages pretrained models to extract image-text pairs from PDFs through layout analysis, OCR, and vision-language pairing, removing the need for manual annotation. Additionally, we construct instruction data from extracted image-text pairs to enrich the training data. To evaluate the effectiveness of PDF-derived data, we train Japanese LMMs and assess their performance on the Japanese LMM Benchmark. Our results demonstrate substantial improvements, with performance gains ranging from 2.1% to 13.8% on Heron-Bench. Further analysis highlights the impact of PDF-derived data on various factors, such as model size and language models, reinforcing its value as a multimodal resource for Japanese LMMs.

多模态日语PDF挖掘自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。