通过时空追踪提升视频文本问答的推理准确性
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
- 引入时空追踪机制恢复视觉实体关系
- 结合OCR线索显著提升答案生成质量
- 适用于需要跨帧理解的视频问答场景
视频文本视觉问答(Video TextVQA)旨在联合推理给定视频中的文本与视觉信息以回答问题。受图像领域TextVQA发展的启发,现有方法采用语言模型(如T5)对多帧富含文本的图像进行处理,并自动回归生成答案。然而,视觉实体(包括场景文字和物体)之间的时空关系会被破坏,模型易受无关信息干扰,导致推理不合理、答案不准确。为此,我们提出TEA(Track the Answer)方法,更有效地将生成式TextVQA框架从图像扩展至视频。TEA以互补方式恢复时空关系,并融入OCR感知线索,增强问题推理质量。在多个公开视频文本VQA数据集上的大量实验验证了该框架的有效性与泛化能力。TEA在性能上大幅超越现有TextVQA方法、视频-语言预训练方法及视频大语言模型。
原文摘要 · Abstract (English)
Video text-based visual question answering (Video TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. Inspired by the development of TextVQA in image domain, existing Video TextVQA approaches leverage a language model (e.g. T5) to process text-rich multiple frames and generate answers auto-regressively. Nevertheless, the spatio-temporal relationships among visual entities (including scene text and objects) will be disrupted and models are susceptible to interference from unrelated information, resulting in irrational reasoning and inaccurate answering. To tackle these challenges, we propose the TEA (stands for ``\textbf{T}rack th\textbf{E} \textbf{A}nswer'') method that better extends the generative TextVQA framework from image to video. TEA recovers the spatio-temporal relationships in a complementary way and incorporates OCR-aware clues to enhance the quality of reasoning questions. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. TEA outperforms existing TextVQA methods, video-language pretraining methods and video large language models by great margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。