arXiv:2410.22736cs.CL2024-10NAACL被引 6

为日语视觉语言模型构建原生多模态数据集,提升性能。

Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model

  • 从网络存档收集日语图文对与交错数据,用现有模型生成指令数据。
  • 基于原生日语数据训练的模型性能优于依赖机器翻译的数据。
  • 适合日语NLP、多模态模型研发人员快速构建本地化数据集。

为开发高性能视觉语言模型(VLM),需准备多模态资源,如图像-文本对、交错数据和指令数据。尽管英语多模态资源丰富,非英语语言如日语却严重缺乏相应数据。为此,本文以日语为例,提出一种从零开始快速构建日语多模态数据集的方法:从网络存档中收集日语图像-文本对与交错数据,并利用现有视觉语言模型直接生成日语指令数据。实验结果表明,基于这些原生数据训练的VLM,在性能上显著优于依赖机器翻译内容的模型。

原文摘要 · Abstract (English)

To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While multimodal resources for English are abundant, there is a significant lack of corresponding resources for non-English languages, such as Japanese. To address this problem, we take Japanese as a non-English language and propose a method for rapidly creating Japanese multimodal datasets from scratch. We collect Japanese image-text pairs and interleaved data from web archives and generate Japanese instruction data directly from images using an existing VLM. Our experimental results show that a VLM trained on these native datasets outperforms those relying on machine-translated content.

多模态日语数据集构建视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。