构建了388万对日语图文数据集,提升AI对日本文化理解能力。
DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
- 通过网页采集+智能过滤+大模型修正,自动化构建高质量数据
- 数据量达388万对,远超现有日语多模态数据集
- 适合研究日本文化场景的视觉语言模型,支持商业应用
针对日语视觉语言建模中高质量大规模资源稀缺的问题,本文提出一种可扩展、可复现的构建流程,融合大规模网络采集、严格过滤去重、基于目标检测的证据提取,以及在接地约束下的大语言模型精修。利用该流程,构建了两个数据集:图像描述数据集DEJIMA-Cap和视觉问答数据集DEJIMA-VQA,各含388万张图像-文本对,显著超过现有日语视觉语言数据集规模。人工评估显示,DEJIMA在日语地道性与语言自然度上优于翻译或人工标注数据,同时事实准确性与人工标注语料相当。图像特征分布的定量分析表明,DEJIMA广泛覆盖日本典型视觉领域,兼具语言与文化代表性。在多个日语多模态基准上,基于DEJIMA训练的模型均表现出持续提升,证实文化语境化、大规模数据对模型性能的关键作用。所有数据源与模块均允许商用,我们公开发布数据集及元数据,以推动日语视觉语言建模的研究与产业应用。
原文摘要 · Abstract (English)
This work addresses the scarcity of high-quality, large-scale resources for Japanese Vision-and-Language (V&L) modeling. We present a scalable and reproducible pipeline that integrates large-scale web collection with rigorous filtering/deduplication, object-detection-driven evidence extraction, and Large Language Model (LLM)-based refinement under grounding constraints. Using this pipeline, we build two resources: an image-caption dataset (DEJIMA-Cap) and a VQA dataset (DEJIMA-VQA), each containing 3.88M image-text pairs, far exceeding the size of existing Japanese V&L datasets. Human evaluations demonstrate that DEJIMA achieves substantially higher Japaneseness and linguistic naturalness than datasets constructed via translation or manual annotation, while maintaining factual correctness at a level comparable to human-annotated corpora. Quantitative analyses of image feature distributions further confirm that DEJIMA broadly covers diverse visual domains characteristic of Japan, complementing its linguistic and cultural representativeness. Models trained on DEJIMA exhibit consistent improvements across multiple Japanese multimodal benchmarks, confirming that culturally grounded, large-scale resources play a key role in enhancing model performance. All data sources and modules in our pipeline are licensed for commercial use, and we publicly release the resulting dataset and metadata to encourage further research and industrial applications in Japanese V&L modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。