评测多模态大模型在视频文字识别中的表现,发现准确率仅73.7%。
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
- 构建涵盖25项任务的MME-VideoOCR基准,覆盖44种视频场景。
- 顶尖模型Gemini-2.5 Pro在视频文字识别中准确率为73.7%。
- 模型在跨帧推理和时空理解上表现不足,适合研究视频理解者关注。
多模态大语言模型(MLLMs)在静态图像文字识别中已取得较高准确率,但在视频场景中因运动模糊、时序变化和视觉特效等因素导致性能显著下降。为更清晰指导实际MLLM训练,我们提出MME-VideoOCR基准,包含10类任务共25个具体任务,覆盖44种多样视频场景。该基准不仅涵盖文字识别,还延伸至对视频中文本内容的深层理解与推理。数据集包含1,464段不同分辨率、宽高比和时长的视频,以及2,000条人工精心标注的问答对。我们在MME-VideoOCR上评估了18个先进MLLMs,结果显示即使最佳模型Gemini-2.5 Pro的准确率也仅为73.7%。细粒度分析表明,现有模型在单帧或少数帧内含相关文本的任务中表现良好,但在需全局视频理解的任务中能力有限,尤其在时空推理、跨帧信息整合及抗语言先验偏见方面表现薄弱。研究还强调高分辨率视觉输入和充分时序覆盖对动态视频中可靠文字识别的重要性。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce the MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1,464 videos with varying resolutions, aspect ratios, and durations, along with 2,000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2.5 Pro) achieves an accuracy of only 73.7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。