arXiv:2510.22276cs.CVcs.CL2025-10

构建首个大规模日文图文数据集,提升视觉语言模型对日本文化的理解能力。

WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

  • 从日本网页爬取1.55亿条原生日文图文对,构建WAON数据集。
  • 在374类日本文化任务上,使用原生数据微调的模型表现优于翻译数据。
  • 适合研究多文化视觉语言模型或日语自然语言处理的研究者使用。

对比学习型视觉语言模型通过大规模预训练取得了显著进展。近期研究表明,去除仅限英文的标题过滤并使用全球数据进行预训练,能有效提升跨文化性能。我们探究此类全球预训练是否足以实现文化特定理解,还是需要进一步利用原生数据进行适应性训练以超越全局预训练的效果。为此,我们提出了WAON——一个基于Common Crawl中日本本土网络内容构建的、目前公开最大的日文图像-文本数据集,包含约1.55亿个样本。同时,我们引入了WAON-Bench,一个涵盖374个类别的手动标注日本文化基准测试集。通过对多个日文图像-文本数据集的对比微调实验发现,使用WAON微调的模型在日文文化基准上的表现持续优于使用英译日数据微调的模型。我们已开源该数据集与代码。

原文摘要 · Abstract (English)

Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing English-only caption filters and pretraining on global data is effective for improving multicultural performance. We study whether such global pretraining is sufficient for culture-specific understanding, or whether further adaptation with natively sourced data can boost performance beyond what global pretraining alone achieves. To enable this investigation, we present WAON, the largest publicly available native Japanese image-text dataset constructed from native Japanese web content in Common Crawl, containing approximately 155 million examples. We also introduce WAON-Bench, a manually curated Japanese cultural benchmark spanning 374 classes. Through comparative fine-tuning experiments on multiple Japanese image-text datasets, we observe that models fine-tuned on WAON consistently achieve stronger performance on Japanese cultural benchmarks than those fine-tuned on English-to-Japanese translated data. We release our dataset and code.

图像文本多语言文化适应数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。