首个综合评估文本生成与编辑能力的基准,涵盖33项真实场景任务。
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
- 统一文本生成、编辑与OCR图像转换,覆盖5类文本与33项任务
- 1060个高密度双语样本,含复杂排版与多尺度文本
- 新指标融合准确率、美观度与指令遵循,揭示模型普遍不足
提升视觉文本合成能力是图像生成模型的长期挑战。尽管当前顶尖模型在文本生成上取得进展,现有基准因范围狭窄(仅限场景文本与海报)、评估孤立(文本生成与编辑分开)且难度不足,难以真实反映性能。为此,我们首次将以文本为中心的文本到图像生成、文本编辑与相关图像到图像的OCR转换统一评估,提出OCRGenBench——迄今最全面的视觉文本合成能力评测基准。该基准涵盖5类常见文本与33项任务,包括文本生成、编辑及文档去畸变、手写消除等。包含1,060个人工标注的指令-图像-真值三元组,刻意设计高文本密度、多尺度、异构比例与双语内容以模拟真实复杂性。我们还引入OCRGenScore,整合文本准确率、美学质量与指令遵循度的统一评分。对19个前沿生成模型的实验表明,多数得分低于60/100。分析揭示了此前被忽视的关键缺陷:文本定位差、内容意外修改,以及对密集或小尺寸文本处理失败。我们希望该基准建立可靠评估标准,推动可信视觉文本合成发展。基准与代码已开源。
原文摘要 · Abstract (English)
Improving visual text synthesis has long been a challenging and evolving frontier for image generation models. While recent state-of-the-art (SOTA) models have made remarkable strides in text generation capabilities, existing benchmarks inadequately assess their true performance due to narrow scope (scene text and posters only), isolated evaluation (T2I generation or editing separately), and insufficient difficulty (lacking challenging scenarios). To bridge this gap, we pioneer the unification of text-centric T2I generation, text editing, and OCR-related image-to-image translation to evaluate a model's holistic visual text synthesis abilities, i.e., OCR generative capabilities. Accordingly, we propose OCRGenBench, the most comprehensive benchmark to date for evaluating these abilities. OCRGenBench covers five common text categories and 33 OCR generative tasks, encompassing T2I generation, text editing, and other image-to-image OCR tasks (e.g., document dewarping and handwriting removal). The benchmark includes 1,060 human-annotated samples consisting of instruction-image-GT triplets, deliberately featuring high text density, diverse generation scales, varied aspect ratios, and bilingual content to capture real-world complexity. Furthermore, we introduce OCRGenScore, a unified metric integrating text accuracy, aesthetic quality, and instruction following. Extensive experiments on 19 cutting-edge generative models reveal that most score below 60/100. Our analysis exposes critical, previously overlooked limitations, including poor text localization, unintended content modifications, and failures with dense or small-scale text. We hope OCRGenBench establishes a robust standard to evaluate OCR generative capabilities, driving the evolution of reliable visual text synthesis. The benchmark and evaluation code are available at https://github.com/NiceRingNode/Awesome-Generative-Models-for-OCR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。