构建首个面向第一人称场景文本的视频问答基准,助力智能助手理解真实场景文字。
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

- 设计1.5千段第一人称视频与7千个场景文本相关问题,贴近真实生活场景。
- 主流大模型最高仅33%准确率,暴露当前技术在动态场景文本理解上的严重不足。
- 强调时间定位、多帧推理和高分辨率输入对提升性能的关键作用,适合视觉语言模型研究者参考。
我们提出EgoTextVQA,一个全新且严谨构建的第一人称场景文本视频问答基准。该数据集包含1.5K段第一人称视角视频和7K个涉及场景文本的问答对,反映户外驾驶与室内家务等真实用户需求。问题设计旨在激发对动态环境中场景文本的识别与推理。我们全面评估了10个主流多模态大语言模型,结果表明所有模型表现不佳,最佳成绩(Gemini 1.5 Pro)仅为33%准确率,凸显现有技术在第一人称问答辅助中的严重缺陷。进一步分析显示,精确的时间定位、多帧推理,以及高分辨率和辅助场景文本输入是提升性能的关键。通过深入分析与启发式建议,我们希望EgoTextVQA能成为第一人称场景文本问答研究的可靠测试平台。数据集已开源:https://github.com/zhousheng97/EgoTextVQA。
原文摘要 · Abstract (English)
We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor driving and indoor house-keeping activities. The questions are designed to elicit identification and reasoning on scene text in an egocentric and dynamic environment. With EgoTextVQA, we comprehensively evaluate 10 prominent multimodal large language models. Currently, all models struggle, and the best results (Gemini 1.5 Pro) are around 33\% accuracy, highlighting the severe deficiency of these techniques in egocentric QA assistance. Our further investigations suggest that precise temporal grounding and multi-frame reasoning, along with high resolution and auxiliary scene-text inputs, are key for better performance. With thorough analyses and heuristic suggestions, we hope EgoTextVQA can serve as a solid testbed for research in egocentric scene-text QA assistance. Our dataset is released at: https://github.com/zhousheng97/EgoTextVQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。