用语言学定义的语义结构提升越南语场景文本理解能力
ViConsFormer: Constituting Meaningful Phrases of Scene Texts using Transformer-based Method in Vietnamese Text-based Visual Question Answering
- 基于语言学语义框架构建文本表示方法
- 在两个越南语数据集上达到当前最佳性能
- 适合需要理解多语言场景文本的应用
文本视觉问答(Text-based VQA)是一项挑战性任务,要求机器利用图像中的场景文本回答给定问题。其核心难点在于从场景文本中提取有意义的信息。现有研究通过嵌入文本边界框的2D坐标来利用空间信息。本文依据语言学中的语义定义,提出一种基于Transformer的新方法,有效挖掘越南语场景文本的信息。实验结果表明,该方法在两个大规模越南语文本视觉问答数据集上均取得当前最优表现。
原文摘要 · Abstract (English)
Text-based VQA is a challenging task that requires machines to use scene texts in given images to yield the most appropriate answer for the given question. The main challenge of text-based VQA is exploiting the meaning and information from scene texts. Recent studies tackled this challenge by considering the spatial information of scene texts in images via embedding 2D coordinates of their bounding boxes. In this study, we follow the definition of meaning from linguistics to introduce a novel method that effectively exploits the information from scene texts written in Vietnamese. Experimental results show that our proposed method obtains state-of-the-art results on two large-scale Vietnamese Text-based VQA datasets. The implementation can be found at this link.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。