arXiv:2410.14991cs.CVcs.CL2024-10中稿 · KDD被引 7

构建首个本土化孟加拉语视觉问答数据集,助力低资源语言AI发展。

ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla

  • 基于本地真实场景构建1.5万+样本的孟加拉语VQA数据集
  • 自研双编码器模型在同类中表现最优,超越主流大模型
  • 专为孟加拉语设计,适合本地化多模态研究与应用

视觉问答(VQA)旨在回答关于视觉内容的自然语言问题。尽管孟加拉语使用广泛,但在VQA领域仍属低资源语言,因缺乏合适基准数据集,且现有数据集缺乏地域相关性,多源自外国版本。为此,我们提出大规模孟加拉语VQA数据集ChitroJera,包含超过1.5万条来自多样化本地数据源的样本。评估了文本编码器、图像编码器、多模态模型及新提出的双编码器模型,实验显示预训练双编码器在同规模模型中表现最佳。同时通过提示工程测试当前大型视觉语言模型(LVLMs),实现整体最优性能。鉴于现有数据集发展滞后,我们期望ChitroJera推动孟加拉语视觉-语言任务的研究拓展。

原文摘要 · Abstract (English)

Visual Question Answer (VQA) poses the problem of answering a natural language question about a visual context. Bangla, despite being a widely spoken language, is considered low-resource in the realm of VQA due to the lack of proper benchmarks, challenging models known to be performant in other languages. Furthermore, existing Bangla VQA datasets offer little regional relevance and are largely adapted from their foreign counterparts. To address these challenges, we introduce a large-scale Bangla VQA dataset, ChitroJera, totaling over 15k samples from diverse and locally relevant data sources. We assess the performance of text encoders, image encoders, multimodal models, and our novel dual-encoder models. The experiments reveal that the pre-trained dual-encoders outperform other models of their scale. We also evaluate the performance of current large vision language models (LVLMs) using prompt-based techniques, achieving the overall best performance. Given the underdeveloped state of existing datasets, we envision ChitroJera expanding the scope of Vision-Language tasks in Bangla.

视觉问答孟加拉语多模态低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。