评测视觉语言模型在动态视频中的文字识别能力,发现其表现优于传统OCR但仍有幻觉等缺陷。
Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
- 构建包含1477帧的开源数据集,覆盖多种动态视频场景。
- VLMs在多数场景下错误率低于传统OCR,最高准确率达92.3%。
- 适合关注视频文字识别与多模态模型评估的研究者使用。
本文提出一个用于评估视觉语言模型(VLMs)在动态视频环境中光学字符识别(OCR)任务的开源基准。研究构建了一个包含1,477帧的手动标注数据集,涵盖代码编辑器、新闻播报、YouTube视频和广告等多种领域。对比了三种前沿VLMs(Claude-3、Gemini-1.5、GPT-4o)与传统OCR系统(EasyOCR、RapidOCR)的表现,评估指标包括词错误率(WER)、字符错误率(CER)和准确率。结果表明,VLMs在多数场景下表现优于传统OCR模型,尤其在语义理解强的文本识别中优势明显。然而,仍存在幻觉、内容安全策略限制以及对遮挡或艺术化字体敏感等问题。数据集与评测框架已公开,以推动该方向研究。
原文摘要 · Abstract (English)
This paper introduces an open-source benchmark for evaluating Vision-Language Models (VLMs) on Optical Character Recognition (OCR) tasks in dynamic video environments. We present a curated dataset containing 1,477 manually annotated frames spanning diverse domains, including code editors, news broadcasts, YouTube videos, and advertisements. Three state of the art VLMs - Claude-3, Gemini-1.5, and GPT-4o are benchmarked against traditional OCR systems such as EasyOCR and RapidOCR. Evaluation metrics include Word Error Rate (WER), Character Error Rate (CER), and Accuracy. Our results highlight the strengths and limitations of VLMs in video-based OCR tasks, demonstrating their potential to outperform conventional OCR models in many scenarios. However, challenges such as hallucinations, content security policies, and sensitivity to occluded or stylized text remain. The dataset and benchmarking framework are publicly available to foster further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。