arXiv:2511.02046cs.CVcs.AI2025-11

用多模型流水线自动生成72000条文本视觉问答数据,省去人工标注

Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis

  • 整合OCR、目标检测、生成模型,构建端到端合成流水线
  • 在4.4万张图像上生成约7.2万对高质量文本视觉问答数据
  • 首个可规模化生成文本视觉问答数据的自动化方案,适合数据集构建者

针对场景文本视觉问答(text-VQA)任务,大规模数据库的创建依赖耗时费力的人工标注。随着视觉语言基础模型和成熟OCR技术的发展,亟需建立一个能基于图像中的场景文本自动生成问题-答案对的端到端流水线。本文提出一种自动化合成text-VQA数据集的流程,可生成真实可靠的问答对,并随场景文本数据量增加而扩展。该方法融合了文本定位(OCR检测与识别)、感兴趣区域(ROI)检测、图像描述生成及问题生成等多个模型与算法,将其整合为统一流水线,实现问答对的自动合成与验证。据我们所知,这是首个能自动合成并验证大规模text-VQA数据集的方案,涵盖约7.2万组问答对,基于约4.4万张图像。

原文摘要 · Abstract (English)

Creation of large-scale databases for Visual Question Answering tasks pertaining to the text data in a scene (text-VQA) involves skilful human annotation, which is tedious and challenging. With the advent of foundation models that handle vision and language modalities, and with the maturity of OCR systems, it is the need of the hour to establish an end-to-end pipeline that can synthesize Question-Answer (QA) pairs based on scene-text from a given image. We propose a pipeline for automated synthesis for text-VQA dataset that can produce faithful QA pairs, and which scales up with the availability of scene text data. Our proposed method harnesses the capabilities of multiple models and algorithms involving OCR detection and recognition (text spotting), region of interest (ROI) detection, caption generation, and question generation. These components are streamlined into a cohesive pipeline to automate the synthesis and validation of QA pairs. To the best of our knowledge, this is the first pipeline proposed to automatically synthesize and validate a large-scale text-VQA dataset comprising around 72K QA pairs based on around 44K images.

文本VQA数据合成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。