arXiv:2603.09689cs.CVcs.AI2026-03

构建大规模越南语视觉问答数据集,推动低资源多模态研究。

AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

  • 基于预训练模型自动构建越南语VQA数据集
  • 验证多种评估指标在多语言场景下的有效性
  • 适合越南语多模态研究与低资源学习方向

视觉问答(VQA)是一项需要模型联合理解视觉与文本信息的基础多模态任务。早期VQA系统过度依赖语言偏见,促使后续研究强调视觉定位与平衡数据集。随着文本与视觉领域大型预训练Transformer的成功,如用于越南语理解的PhoBERT和用于图像表征学习的Vision Transformers(ViT),多模态融合取得显著进展。针对越南语VQA,已有若干数据集被提出以促进低资源多模态学习,包括ViVQA、OpenViVQA及最近发布的ViTextVQA,这些资源支持在越南语背景下整合语言与视觉特征的模型基准测试。现有VQA系统评估常使用原本为图像描述或机器翻译设计的自动指标,如BLEU、METEOR、CIDEr、Recall、Precision和F1-score。然而近期研究表明,大语言模型可进一步提升自动评估与人类判断的一致性。本文探索基于Transformer架构的越南语视觉问答,结合文本与视觉预训练,并在多语言设置下系统比较自动评估指标的有效性。

原文摘要 · Abstract (English)

Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA systems relied heavily on language biases, motivating subsequent work to emphasize visual grounding and balanced datasets. With the success of large-scale pre-trained transformers for both text and vision domains -- such as PhoBERT for Vietnamese language understanding and Vision Transformers (ViT) for image representation learning -- multimodal fusion has achieved remarkable progress. For Vietnamese VQA, several datasets have been introduced to promote research in low-resource multimodal learning, including ViVQA, OpenViVQA, and the recently proposed ViTextVQA. These resources enable benchmarking of models that integrate linguistic and visual features in the Vietnamese context. Evaluation of VQA systems often employs automatic metrics originally designed for image captioning or machine translation, such as BLEU, METEOR, CIDEr, Recall, Precision, and F1-score. However, recent research suggests that large language models can further improve the alignment between automatic evaluation and human judgment in VQA tasks. In this work, we explore Vietnamese Visual Question Answering using transformer-based architectures, leveraging both textual and visual pre-training while systematically comparing automatic evaluation metrics under multilingual settings.

视觉问答越南语多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。