首个视频OCR评测基准,测试多模态模型识字能力。
Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study
- 构建包含1028段视频的自动化评测集
- 涵盖文本识别、语义理解、动态定位6类任务
- 适合研究视频大模型文本理解的研究者使用
随着多模态大语言模型的兴起,从视频内容中准确提取和理解文本信息(即基于视频的光学字符识别,Video OCR)已成为关键能力。本文提出一个新型基准,用于评估多模态模型在视频中的OCR性能。该基准包含1,028个视频和2,961个问答对,通过6个不同子任务设计了若干关键挑战:(1) 文本内容及其基本视觉属性的识别;(2) 视频中OCR对象的语义与空间理解;(3) 动态运动检测与时间定位。我们采用半自动化方法构建该基准,结合图像LLM的OCR能力与人工精修,在效率、成本与数据质量间取得平衡。本资源旨在推动视频大模型研究,并凸显提升视频大模型OCR能力的迫切需求。基准将发布于https://github.com/YuHuiGao/FG-Bench.git。
原文摘要 · Abstract (English)
With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character recognition (Video OCR), has become a crucial capability. This paper introduces a novel benchmark designed to evaluate the video OCR performance of multi-modal models in videos. Comprising 1,028 videos and 2,961 question-answer pairs, this benchmark proposes several key challenges through 6 distinct subtasks: (1) Recognition of text content itself and its basic visual attributes, (2)Semantic and Spatial Comprehension of OCR objects in videos (3) Dynamic Motion detection and Temporal Localization. We developed this benchmark using a semi-automated approach that integrates the OCR ability of image LLMs with manual refinement, balancing efficiency, cost, and data quality. Our resource aims to help advance research in video LLMs and underscores the need for improving OCR ability for video LLMs. The benchmark will be released on https://github.com/YuHuiGao/FG-Bench.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。