通过扫描聚焦放大,让大模型更准地回答视频文字问题。
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
- 模拟人类答题流程,三步引导注意力:扫描、聚焦、放大关键文本
- 在多个公开数据集上达到新最好效果,显著超越此前方法
- 无需训练,直接增强Video-LLM能力,适合快速部署到各类视频问答场景
视频文本视觉问答(Video TextVQA)任务旨在利用视频中出现的视觉文本回答相关问题。该任务面临严峻挑战:需准确感知和理解跨帧尺度、方向与清晰度各异的场景文字,并有效融合时序与语义上下文以生成精确答案。同时,模型必须识别与问题相关的文字线索,过滤冗余信息,确保回答由最相关、最有信息量的线索引导。为此,我们提出SFA——首个面向Video TextVQA的基于Video-LLM且无需训练的框架,灵感源自人类答题过程。通过自适应扫描视频帧、选择性聚焦关键区域并直接放大,SFA有效引导Video-LLM注意力至核心线索,从而提升答案准确性。SFA在多个公开Video TextVQA数据集上取得新最优性能,显著超越先前方法,验证了其有效性与泛化能力。
原文摘要 · Abstract (English)
Video text-based visual question answering (Video TextVQA) task aims to answer questions about videos by leveraging the visual text appearing within the videos. This task poses significant challenges, requiring models to accurately perceive and comprehend scene text that varies in scale, orientation, and clarity across frames, while effectively integrating temporal and semantic context to generate precise answers. Moreover, the model must identify question-relevant textual cues and filter out redundant or irrelevant information to ensure answering is guided by the most relevant and informative cues. To address these challenges, we propose SFA, a training-free framework and the first Video-LLM-based method tailored for Video TextVQA, motivated by the human process of answering questions. By adaptively scanning video frames, selectively focusing on key regions, and directly amplifying them, SFA effectively guides the Video-LLM's attention toward essential cues, enabling it to generate more accurate answers. SFA achieves new state-of-the-art results across several public Video TextVQA datasets and surpasses previous methods by a substantial margin, demonstrating its effectiveness and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。