首个越南语路牌视觉问答数据集,提升低资源语言文本理解能力
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
- 构建越南语路牌VQA数据集,融合多语言、文化与视觉特征
- OCR增强上下文使准确率最高提升209%
- 多智能体框架结合GPT-4实现75.98%准确率,适合跨模态研究者
理解自然场景中的路牌文本对视觉问答(VQA)的实际应用至关重要,但在低资源语言中仍研究不足。我们提出ViSignVQA,首个大规模越南语路牌导向VQA数据集,包含10,762张图像和25,573个问答对。数据集涵盖越南路牌的多语言、文化与视觉特征,如双语文字、非正式表达及颜色、版式等视觉元素。为评估该任务,我们集成越南语OCR模型(SwinTextSpotter)与预训练语言模型(ViT5),适配BLIP-2、LaTr、PreSTU和SaL等先进VQA模型。实验表明,引入OCR文本可使F1分数最高提升209%。此外,我们提出一种结合感知与推理智能体的多智能体框架,利用GPT-4多数投票,达到75.98%准确率。本研究首次建立大规模越南语路牌理解多模态数据集,凸显领域专用资源在低资源语言文本VQA中的重要性。ViSignVQA作为基准,捕捉真实场景文本特征,支持越南语OCR集成VQA模型的开发与评估。
原文摘要 · Abstract (English)
Understanding signboard text in natural scenes is essential for real-world applications of Visual Question Answering (VQA), yet remains underexplored, particularly in low-resource languages. We introduce ViSignVQA, the first large-scale Vietnamese dataset designed for signboard-oriented VQA, which comprises 10,762 images and 25,573 question-answer pairs. The dataset captures the diverse linguistic, cultural, and visual characteristics of Vietnamese signboards, including bilingual text, informal phrasing, and visual elements such as color and layout. To benchmark this task, we adapted state-of-the-art VQA models (e.g., BLIP-2, LaTr, PreSTU, and SaL) by integrating a Vietnamese OCR model (SwinTextSpotter) and a Vietnamese pretrained language model (ViT5). The experimental results highlight the significant role of the OCR-enhanced context, with F1-score improvements of up to 209% when the OCR text is appended to questions. Additionally, we propose a multi-agent VQA framework combining perception and reasoning agents with GPT-4, achieving 75.98% accuracy via majority voting. Our study presents the first large-scale multimodal dataset for Vietnamese signboard understanding. This underscores the importance of domain-specific resources in enhancing text-based VQA for low-resource languages. ViSignVQA serves as a benchmark capturing real-world scene text characteristics and supporting the development and evaluation of OCR-integrated VQA models in Vietnamese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。