arXiv:2509.09014cs.CVcs.CL2025-09

构建首个大规模乌尔都图文数据集,提升低资源语言视觉理解能力

COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

  • 基于分层采样构建5.9万图像、31.9万句乌尔都语图文对
  • 融合多模态评估框架自动筛选高质量翻译,迭代优化文本
  • 开源数据集与质检流程,助力公平多元的多模态研究

乌尔都语是超过2.5亿人的母语,但在多模态与视觉-语言研究中仍严重缺乏支持。现有大型高质量数据集的缺失限制了乌尔都语系统的发展,并加剧了以高资源语言训练的多语言视觉-语言模型中的偏见。为此,我们提出了COCO-Urdu,一个源自MS COCO的大规模图像-标题数据集,包含59,000张图像和319,000条乌尔都语标题,通过分层采样保留原始分布。标题使用SeamlessM4T v2翻译,并通过混合多模态质量评估框架进行验证:结合COMET-Kiwi评估翻译质量,基于CLIP的相似度衡量视觉关联性,以及BERTScore与回译法确保语义一致性;低分标题通过开源大语言模型迭代修正。我们在BLEU、SacreBLEU和chrF上进行了基准测试,结果表现一致优异。据我们所知,COCO-Urdu是目前公开可用的最大乌尔都语图文数据集。通过发布数据集与质量评估流水线,我们旨在减少多模态研究中的语言偏见,为包容性视觉-语言系统奠定基础。

原文摘要 · Abstract (English)

Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. The absence of large-scale, high-quality datasets has limited the development of Urdu-capable systems and reinforced biases in multilingual vision-language models trained primarily on high-resource languages. To address this gap, we present COCO-Urdu, a large-scale image-caption dataset derived from MS COCO, containing 59,000 images and 319,000 Urdu captions selected through stratified sampling to preserve the original distribution. Captions were translated using SeamlessM4T v2 and validated with a hybrid multimodal quality estimation framework that integrates COMET-Kiwi for translation quality, CLIP-based similarity for visual grounding, and BERTScore with back-translation for semantic consistency; low-scoring captions were iteratively refined using open-source large language models. We further benchmark COCO-Urdu on BLEU, SacreBLEU, and chrF, reporting consistently strong results. To the best of our knowledge, COCO-Urdu is the largest publicly available Urdu captioning dataset. By releasing both the dataset and the quality estimation pipeline, we aim to reduce language bias in multimodal research and establish a foundation for inclusive vision-language systems.

图文生成多模态低资源语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。