系统梳理视觉文本处理进展,提出新评估框架和基准。
Visual Text Processing: A Comprehensive Review and Unified Evaluation
- 从多角度分析文本特征与任务匹配关系。
- 构建涵盖多种任务的VTPBench基准,覆盖20+模型实验。
- 引入MLLM驱动的VTPScore评估指标,提升评测可靠性。
视觉文本在文档与场景图像中承载丰富语义信息,是计算机视觉领域的研究重点。随着基础模型的发展,该领域已超越传统文字检测与识别任务,拓展至图像重建与编辑等新方向。然而,文本的独特属性仍带来挑战,有效捕捉其特征对构建鲁棒模型至关重要。本文系统回顾了视觉文本处理的最新进展,聚焦两个核心问题:(1) 不同任务适配的文本特征类型;(2) 如何有效融入处理框架。我们提出VTPBench基准,整合多种视觉文本数据集,并利用多模态大语言模型(MLLMs)的视觉质量评估能力,设计新型评估指标VTPScore,确保评价公平可靠。基于超过20个具体模型的实证研究显示,现有技术仍有显著提升空间。本工作旨在为该动态领域提供基础资源,推动未来创新。相关代码库见:https://github.com/shuyansy/Visual-Text-Processing-survey。
原文摘要 · Abstract (English)
Visual text is a crucial component in both document and scene images, conveying rich semantic information and attracting significant attention in the computer vision community. Beyond traditional tasks such as text detection and recognition, visual text processing has witnessed rapid advancements driven by the emergence of foundation models, including text image reconstruction and text image manipulation. Despite significant progress, challenges remain due to the unique properties that differentiate text from general objects. Effectively capturing and leveraging these distinct textual characteristics is essential for developing robust visual text processing models. In this survey, we present a comprehensive, multi-perspective analysis of recent advancements in visual text processing, focusing on two key questions: (1) What textual features are most suitable for different visual text processing tasks? (2) How can these distinctive text features be effectively incorporated into processing frameworks? Furthermore, we introduce VTPBench, a new benchmark that encompasses a broad range of visual text processing datasets. Leveraging the advanced visual quality assessment capabilities of multimodal large language models (MLLMs), we propose VTPScore, a novel evaluation metric designed to ensure fair and reliable evaluation. Our empirical study with more than 20 specific models reveals substantial room for improvement in the current techniques. Our aim is to establish this work as a fundamental resource that fosters future exploration and innovation in the dynamic field of visual text processing. The relevant repository is available at https://github.com/shuyansy/Visual-Text-Processing-survey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。