通过动态课程学习提升越南语视觉问答性能
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
- 用改写增强文本与动态课程学习结合,分阶段训练模型
- 在OpenViVQA上准确率显著提升,ViVQA表现波动但具潜力
- 适合资源匮乏语言的视觉问答研究,尤其关注越南语
视觉问答(VQA)是需要跨文本与视觉输入进行推理的多模态任务,在越南语等低资源语言中因语言多样性及高质量数据集缺失而尤为困难。传统方法依赖大量标注数据、高计算开销的流程和大型预训练模型,限制了其在越南语VQA中的应用。为此,我们提出一种结合基于改写的特征增强模块与动态课程学习策略的训练框架。具体而言,增强样本视为“简单”样本,原始样本视为“困难”样本,框架通过动态调整难易样本比例,逐步提高同一数据集的难度。该机制促进模型渐进适应任务复杂性,从而提升泛化能力。实验表明,该方法在OpenViVQA数据集上实现持续性能提升,在ViVQA数据集上结果混合,凸显了其在推进越南语VQA方面的潜力与挑战。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) is a multimodal task requiring reasoning across textual and visual inputs, which becomes particularly challenging in low-resource languages like Vietnamese due to linguistic variability and the lack of high-quality datasets. Traditional methods often rely heavily on extensive annotated datasets, computationally expensive pipelines, and large pre-trained models, specifically in the domain of Vietnamese VQA, limiting their applicability in such scenarios. To address these limitations, we propose a training framework that combines a paraphrase-based feature augmentation module with a dynamic curriculum learning strategy. Explicitly, augmented samples are considered "easy" while raw samples are regarded as "hard". The framework then utilizes a mechanism that dynamically adjusts the ratio of easy to hard samples during training, progressively modifying the same dataset to increase its difficulty level. By enabling gradual adaptation to task complexity, this approach helps the Vietnamese VQA model generalize well, thus improving overall performance. Experimental results show consistent improvements on the OpenViVQA dataset and mixed outcomes on the ViVQA dataset, highlighting both the potential and challenges of our approach in advancing VQA for Vietnamese language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。