arXiv:2508.19887cs.CLcs.CV2025-08被引 5

构建5.2万对孟加拉语视觉问答数据集,提升低资源语言AI研究质量

Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement

  • 用多语言大模型辅助翻译优化,减少人工误差
  • 含52,650组问答对,覆盖3类答案类型和4750+图像
  • 为低资源语言多模态研究提供高质量开源基准

本文提出 Bangla-Bayanno,一个面向孟加拉语的开放性视觉问答(VQA)数据集,该语言在多模态人工智能研究中属广泛使用但资源稀缺的语言。现有数据集多为特定领域、问题类型或答案格式的人工标注,或受限于小众答案形式。为降低人为错误并确保表述清晰,我们采用多语言大模型辅助的翻译精炼流程,克服了多语言源翻译质量不佳的问题。该数据集包含超过4750张图像上的52,650个问答对,问题按答案类型分为三类:名词型(简短描述)、数量型(数值)和是非型(是/否)。Bangla-Bayanno 是目前最全面、高质量的开源孟加拉语VQA基准,旨在推动低资源多模态学习研究,并促进更包容的AI系统发展。

原文摘要 · Abstract (English)

In this paper, we introduce Bangla-Bayanno, an open-ended Visual Question Answering (VQA) Dataset in Bangla, a widely used, low-resource language in multimodal AI research. The majority of existing datasets are either manually annotated with an emphasis on a specific domain, query type, or answer type or are constrained by niche answer formats. In order to mitigate human-induced errors and guarantee lucidity, we implemented a multilingual LLM-assisted translation refinement pipeline. This dataset overcomes the issues of low-quality translations from multilingual sources. The dataset comprises 52,650 question-answer pairs across 4750+ images. Questions are classified into three distinct answer types: nominal (short descriptive), quantitative (numeric), and polar (yes/no). Bangla-Bayanno provides the most comprehensive open-source, high-quality VQA benchmark in Bangla, aiming to advance research in low-resource multimodal learning and facilitate the development of more inclusive AI systems.

视觉问答低资源语言多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。