arXiv:2503.18746cs.CV2025-03CVPR被引 12

让图像模型学会用语言理解文字,提升模糊场景下的识别准确率。

Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition

  • 设计语言感知模块,从不同视觉外观中提取独立于图像的语言特征。
  • 在多个基准上达到顶尖性能,显著优于现有自监督方法。
  • 适合需要低标注成本的文本识别场景,尤其对模糊图像效果好。

文本图像兼具视觉与语言双重特性:视觉部分包含结构和外观特征,语言部分包含上下文与语义信息。在视觉质量退化的场景下,语言模式对理解至关重要,因此需融合两者实现鲁棒的场景文本识别(STR)。当前方法多依赖语言模型或语义推理模块,通常需大规模标注数据。自监督学习因无标注,难以分离全局语境相关的语言特征。序列对比学习侧重局部特征对齐,而掩码图像建模(MIM)则聚焦局部结构重建,导致语言知识利用不足。本文提出语言感知掩码图像建模(LMIM),通过独立分支将语言信息引入MIM解码过程。具体设计语言对齐模块,利用不同视觉外观输入提取与视觉无关的语言引导特征。由于特征超越局部结构,LMIM需考虑全局上下文以实现重建。大量实验表明,该方法在多个基准上达到领先水平,注意力可视化也证实其同时捕捉了视觉与语言信息。

原文摘要 · Abstract (English)

Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, linguistic patterns serve as crucial supplements for comprehension, highlighting the necessity of integrating both aspects for robust scene text recognition (STR). Contemporary STR approaches often use language models or semantic reasoning modules to capture linguistic features, typically requiring large-scale annotated datasets. Self-supervised learning, which lacks annotations, presents challenges in disentangling linguistic features related to the global context. Typically, sequence contrastive learning emphasizes the alignment of local features, while masked image modeling (MIM) tends to exploit local structures to reconstruct visual patterns, resulting in limited linguistic knowledge. In this paper, we propose a Linguistics-aware Masked Image Modeling (LMIM) approach, which channels the linguistic information into the decoding process of MIM through a separate branch. Specifically, we design a linguistics alignment module to extract vision-independent features as linguistic guidance using inputs with different visual appearances. As features extend beyond mere visual structures, LMIM must consider the global context to achieve reconstruction. Extensive experiments on various benchmarks quantitatively demonstrate our state-of-the-art performance, and attention visualizations qualitatively show the simultaneous capture of both visual and linguistic information.

文本识别自监督语言建模图像建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。